xAI releases Grok-1.5V, a multimodal AI model with vision support
xAI, an artificial intelligence company under Musk, announced the launch of its first multimodal AI model, Grok-1.5V. In addition to its powerful text processing capabilities, Grok can also process various visual information, including documents, charts, screenshots, and photos. In benchmark tests in multiple fields, Grok-1.5V's performance is comparable to existing cutting-edge multimodal models. Especially in xAI's newly launched RealWorldQA benchmark test, Grok surpassed similar models in its ability to understand the real-world space. The RealWorldQA dataset contains more than 700 images and aims to evaluate the basic understanding ability of multimodal models in the physical world. Grok-1.5 will soon be open to early testers and existing users.
Disclaimer: The content of this article solely reflects the author's opinion and does not represent the platform in any capacity. This article is not intended to serve as a reference for making investment decisions.
You may also like
[Initial listing] Bitget to list Theoriq (THQ). Grab a share of 3,016,600 THQ
CandyBomb x VSN: Trade VSN, XRP or SOL to share 2,931,200 VSN
New users get a 100 USDT margin gift—Trade to earn up to 1088 USDT!
Subscribe to ETH Earn products for dual rewards exclusive for VIPs— Enjoy up to 10% APR and trade to unlock an additional pool of 50,000 USDT
