Source: Electronic Times, reporter Lee Won-ji
Full quotation from the article “CLEVI ranks in the top 2.5% of 3,090 models with 79.07% on GAIA using its own from scratch model”
Publication date: 2026-04-07
Link: https://www.etnews.com/20260407000203
cip-5.5-agent·cip-5.5-mm proprietary model stack: all five agents score above 70
98%+ accuracy when accounting for information loss, surpassing humans (92%)
AI start-up 'CLEVI' recorded an accuracy rate of 79.07% on the global AI agent benchmark 'GAIA' using its own from scratch model, placing it within the top 2.5% of all 3,090 registered models.
According to industry analysts, GAIA is regarded as accurately reflecting the capabilities of 'practical AI agents' in the AI benchmark market. Jointly developed by Meta AI, HuggingFace and Université Paris-Saclay, this benchmark was first released in November 2023.
GAIA is distinguished from other benchmarks in three key ways. First, overfitting is impossible because the answers are not disclosed. Second, high scores cannot be achieved through simple pattern matching because it requires multistep and multimodal complex reasoning. Third, more than 3,090 models from around the world compete under identical conditions via the official HuggingFace leaderboard.
The test set consists of 301 questions in total. Level 1 involves using a single tool and simple reasoning, Level 2 involves combining multiple tools and intermediate-level reasoning, and Level 3 consists of the highest-difficulty questions requiring agent planning and execution across five or more steps. At the time the benchmark was announced, the score was only around 15% for GPT models with plugins, while leading teams have now raised it to the 70–80% range.
Among these results, CLEVI emphasised that the most technically noteworthy aspect was that it did not use any external LLM APIs, such as GPT, Claude or Gemini. A key feature is that CLEVI used two models developed independently from scratch.
cip-5.5-agent is an agentic AI model that serves as the core brain of an agent pipeline that autonomously plans, executes and verifies complex tasks. cip-5.5-mm is a high-performance general-purpose multimodal model that understands and reasons over various file formats, including voice, images and video. As a significant number of GAIA questions include non-text inputs such as PDFs, image files and audio files, this multimodal capability became a key factor in achieving a high score.
Based on these two proprietary models, CLEVI entered GAIA with five different agent configurations, all of which scored above 70.
According to the company, an interesting result was identified during CLEVI’s internal post-review. For some of the 301 questions in total, the evidence supporting the correct answers had disappeared from the publicly available web. In some cases, information that was previously available online or through search engines had been deleted or changed and was no longer accessible. When the questions were reassessed based on the information currently publicly verifiable by humans, CLEVI’s accuracy was found to be above 98%. This figure already exceeds the human average of 92%.
A CLEVI representative explained: “CLEVI recorded scores of 70+ for all five agents and entered the top 2.5% in the public benchmark using only its own models developed from scratch and its own AI agent solutions, without mixing external models. Further improvements to its official score are expected through future enhancements to its own models and agent solutions. This achievement is significant because it demonstrates ‘validated trust’ measured on a global public benchmark; represents an achievement based on a from-scratch proprietary model by a domestic AI developer; is the result of an independent model stack rather than a mixture of external models; and, considering the loss of information for some questions, is meaningful beyond the score itself.”
