Source: ETNews reporter Lee Won-ji
Full article quotation: “CLEVI, 79.07% on GAIA with its own from scratch model... Top 2.5% among 3,090 models”
Publication date: 2026-04-07
Link: https://www.etnews.com/20260407000203
cip-5.5-agent·cip-5.5-mm proprietary model stack; all five agents scored in the 70s or higher
98%+ accuracy when accounting for information loss, surpassing humans (92%)
AI startup 'CLEVI' recorded an accuracy rate of 79.07% on the global AI agent benchmark 'GAIA' using its own from scratch model, placing it within the top 2.5% among all 3,090 registered models.
According to industry experts, GAIA is regarded in the AI benchmark market as accurately reflecting the capabilities of 'practical AI agents'. Jointly developed by Meta AI, HuggingFace, and Université Paris-Saclay, this benchmark was first released in November 2023.
GAIA is distinguished from other benchmarks in three key ways. First, overfitting is impossible because the answers are not disclosed. Second, high scores cannot be achieved through simple pattern matching because it requires multistep, multimodal complex reasoning. Third, more than 3,090 models from around the world compete under identical conditions through the official Hugging Face leaderboard.
The test set consists of a total of 301 questions. Level 1 involves using a single tool and simple reasoning, Level 2 involves combining multiple tools and intermediate-level reasoning, and Level 3 consists of the most difficult questions, requiring agent planning and execution across five or more steps. At the time the benchmark was announced, the score was approximately 15% for GPT models (with plugins), while leading teams have now raised it to the 70–80% range.
Among these results, CLEVI emphasized that the most technically notable aspect was that it did not use any external LLM APIs, such as GPT, Claude, or Gemini. A distinctive feature of CLEVI is its use of two models independently developed from scratch.
cip-5.5-agent is an agentic AI model that serves as the core brain of agent pipelines that autonomously plan, execute and verify complex tasks. cip-5.5-mm is a high-performance general-purpose multimodal model that understands and reasons over various file formats, including audio, images and video. Since a significant number of GAIA questions include non-text inputs such as PDF, image and audio files, this multimodal capability became a key factor in achieving a high score.
Based on these two proprietary models, CLEVI entered GAIA with five different agent configurations, all of which scored 70 or higher.
According to the company, an interesting result was identified during CLEVI’s internal post-review. For some of the 301 questions in total, the supporting evidence for the correct answers had disappeared from the publicly available web. In some cases, information that had previously been available online or through search engines had been deleted or changed and was no longer accessible. When the questions were reevaluated based on the information currently available for public verification, CLEVI’s accuracy was found to be over 98%. This already exceeds the human average of 92%.
A CLEVI representative explained, “Without mixing external models, CLEVI recorded scores of 70+ for all five agents using only its from-scratch proprietary models and proprietary AI agent solutions, placing it in the top 2.5% of the public benchmark. Further improvements to its proprietary models and agent solutions are expected to enable additional improvements to its official score in the future. This achievement is particularly significant because it demonstrates ‘verified trust’ measured on a global public benchmark; represents results based on from-scratch proprietary models developed by a domestic AI company; reflects the outcome of an independent model stack rather than a mixture of external models; and, considering the loss of information for some questions, merits interpretation beyond its score.”
