Source: Electronic Times, Reporter Lee Won-ji
Full article citation: “CLEVI achieves 79.07% on GAIA with its own from scratch model... Among the top 2.5% of 3,090 models”
Publication date: 2026-04-07
Link: https://www.etnews.com/20260407000203
cip-5.5-agent·cip-5.5-mm proprietary model stack: all five agents scored above 70
98%+ accuracy when accounting for information loss, surpassing humans (92%)
AI startup 'CLEVI' recorded an accuracy of 79.07% on the global AI agent benchmark 'GAIA' using its own from scratch model, placing it within the top 2.5% based on all 3,090 registered models.
According to industry commentary, GAIA is regarded in the AI benchmark market as accurately reflecting the capabilities of 'practical AI agents'. Jointly developed by Meta AI, HuggingFace and Université Paris-Saclay, this benchmark was first unveiled in November 2023.
GAIA is distinguished from other benchmarks in three key ways. First, overfitting is impossible because the answers are private. Second, high scores cannot be achieved through simple pattern matching because it requires multi-step, multimodal complex reasoning. Third, more than 3,090 models worldwide compete under identical conditions through the official HuggingFace leaderboard.
The test set consists of 301 questions in total. Level 1 involves using a single tool and simple reasoning, Level 2 involves combining multiple tools and intermediate-level reasoning, and Level 3 consists of the highest-difficulty questions requiring agent planning and execution across five or more steps. At the time the benchmark was announced, GPT models (with plugins) achieved only around 15%, while leading teams have now raised this to the 70–80% range.
Among these results, CLEVI emphasised that the most technically noteworthy aspect was that it did not use external LLM APIs such as GPT, Claude or Gemini at all. A distinctive feature is that CLEVI used two independently developed models created using the from scratch approach.
cip-5.5-agent is an agentic AI model that serves as the core brain of agent pipelines that autonomously plan, execute and verify complex tasks. cip-5.5-mm is a high-performance general-purpose multimodal model that understands and reasons over various file formats, including audio, images and video. Since a significant number of GAIA questions include non-text inputs such as PDFs, image files and audio files, this multimodal capability became a key factor in achieving a high score.
CLEVI entered GAIA with five different agent configurations based on these two proprietary models, and all of them scored above 70.
According to the company, an interesting result was identified in CLEVI’s internal post-evaluation review. For some of the 301 questions in total, the sources supporting the correct answers had disappeared from the publicly available web. In these cases, information that was previously available online or through search engines had been deleted or changed and was no longer accessible. When reevaluated based on the information currently verifiable by people in the public domain, CLEVI’s accuracy was found to be above 98%. This figure already exceeds the human average of 92%.
A CLEVI representative explained, “Without mixing external models, CLEVI achieved scores of 70+ across all five agents using only its own from scratch models and proprietary AI agent solutions, placing it in the top 2.5% of the public benchmark. Further improvements to its proprietary models and agent solutions are expected to enable additional gains in the official score. This achievement is particularly significant because it demonstrates ‘verified trust’ measured on a global public benchmark; represents an achievement by a domestic AI developer based on from scratch proprietary models; is the result of an independent model stack rather than a mix of external models; and, considering the loss of information for some questions, has interpretive value beyond the score itself.”
