Source: Digital Times, reporter Gu Bon-gyu
Full article quoted: “AI start-up Clevi enters the top 2.5% of GAIA… demonstrating proven credibility”
Publication date: 8 April 2026
Link: https://n.news.naver.com/article/029/0003020511?sid=105
South Korean AI start-up ‘Clevi’ has built its own model and agent solution from scratch, and announced that it had become the first in South Korea to enter the top 2.5% of the GAIA benchmark, based on all 3,090 registered models.
According to industry sources, GAIA, designed by Meta AI and operated by Hugging Face, tests AI not on its ‘ability to do one thing well’, but on its ‘ability to combine multiple capabilities to solve real-world problems’. It must find information on the web, read tables in PDFs, analyse images, run code to perform calculations, and synthesise the results to produce a single correct answer. With 301 private questions, three levels of difficulty, and a policy of keeping answers confidential, it is structured to prevent both overfitting and cheating.
Two things are required to achieve a high score on this test. The first is the reasoning ability of the underlying language model, and the second is an agent solution that connects the model to real-world tools and enables it to perform tasks autonomously. No matter how good the model is, if the agent is weak, it cannot solve Level 3; and no matter how sophisticated the agent is, if the model’s reasoning ability is insufficient, incorrect answers will propagate from the intermediate stages.
The current structure of the GAIA leaderboard reveals an interesting pattern. Most top-ranking agents have adopted a ‘multi-model mix’ strategy that combines multiple Big Tech models, such as GPT, Claude and Gemini. This is a rational approach that combines the strengths of each model and protects against score reductions caused by information loss and other factors.
Clevi, however, has not joined that group. Developed from scratch, cip-5.5-agent (agentic AI) autonomously plans, executes and verifies complex tasks, while cip-5.5-mm (high-performance general-purpose multimodal) understands and reasons across various file formats, including audio, images and video. Both models were independently developed within Clevi, and no external LLM APIs were used at all.
And with this proprietary model stack, it entered five agents in GAIA, all of which scored above 70. From a top score of 79.07% to 70.76%, this demonstrated the stability of the entire stack rather than the luck of a single result.
According to the company, there is another figure behind the official score of 79.07%. CLEVI’s internal post-review found that a considerable number of the 301 questions no longer had supporting evidence for the correct answers available on the public web. When reassessed using only questions for which the correct answer still exists online, CLEVI’s accuracy was above 98%. This is already higher than the human average of 92%.
Of course, the official score could have been raised further by combining powerful general-purpose models from other companies. However, CLEVI deliberately chose not to adopt that strategy. This is because the goal of this benchmark was not to compete for the highest score, but to verify whether a from scratch proprietary model had reached a level capable of competing on the global stage.
Having its name listed on the GAIA leaderboard is highly significant because it provides verified credibility through an independent third-party evaluation. This serves as objective evidence that can be used across all stages, including fundraising, international expansion and B2B sales.
A CLEVI representative said, “Many of the 63 officially incorrect answers resulted from output format mismatches, errors in interpreting visual materials, and questions that could not be answered due to information loss. These are issues with post-processing precision and the availability of external information, rather than fundamental limitations in logical reasoning. Because it is a proprietary model, CLEVI can improve it itself. This is precisely the structural advantage of a from scratch proprietary model. CLEVI’s achievement this time poses the question to Korea’s AI industry: ‘How long will we continue to depend on borrowed brains?’ CLEVI has chosen the more difficult path of from scratch development and aims to prove the answer on the global stage,” he said.
