News

AI start-up Clevi enters the top 2.5% of GAIA… demonstrating proven credibility

Source: Digital Times, reporter Bon-gyu Gu

Source: Digital Times, reporter Bon-gyu Gu

Full quotation from the article “AI start-up Clevi enters the top 2.5% of GAIA… demonstrating proven credibility”

Publication date: 8 April 2026

Link: https://n.news.naver.com/article/029/0003020511?sid=105

South Korean AI start-up ‘Clevi’ announced that, while building its own model and proprietary agent solution developed from scratch, it had become the first in South Korea to enter the top 2.5% of the GAIA benchmark, based on all 3,090 registered models.

According to industry sources, GAIA, designed by Meta AI and operated by Hugging Face, tests AI not on its ‘ability to do one thing well’, but on its ‘ability to combine multiple capabilities to solve real-world problems’. It must find information on the web, read tables in PDFs, analyse images, run code to perform calculations, and synthesise the results to provide a single correct answer. With 301 confidential questions, three levels of difficulty, and a policy of not disclosing the answers, it is structured to make both overfitting and cheating impossible.

Two things are required to achieve a high score on this test. The first is the reasoning ability of the underlying language model, and the second is an agent solution that connects the model to real-world tools and enables it to perform tasks autonomously. No matter how good the model is, if the agent is weak, it cannot solve Level 3; and no matter how sophisticated the agent is, if the model’s reasoning ability is insufficient, incorrect answers will be propagated at the intermediate stages.

The current structure of the GAIA leaderboard reveals an interesting pattern. Most of the leading agents have adopted a ‘multi-model mix’ strategy that combines several Big Tech models, such as GPT, Claude and Gemini. This is a rational approach that combines the strengths of each model and guards against score reductions caused by information loss and other factors.

Clevi, however, has not joined that group. Developed from scratch, cip-5.5-agent (agentic AI) autonomously performs the planning, execution and verification of complex tasks, while cip-5.5-mm (high-performance general-purpose multimodal) understands and reasons across various file formats, including audio, images and video. Both models were independently developed internally by Clevi, and no external LLM APIs were used at all.

And by entering five agents built on this proprietary model stack in GAIA, all of them recorded scores in the 70s or above. From the highest score of 79.07% to 70.76%, this demonstrated the stability of the entire stack, rather than the luck of a single result.

Main text image

According to the company, another figure lies behind the official score of 79.07%. CLEVI’s internal post-review found that, among the 301 questions, a considerable number had lost their supporting evidence for the correct answers on the publicly available web. When reassessed based only on questions for which the correct answer still exists on the web, CLEVI’s accuracy rate was over 98%. This figure has already surpassed the human average of 92%.

Of course, mixing in powerful general-purpose models from other companies could have raised the official score further. However, CLEVI deliberately did not adopt that strategy. This was because the goal of this benchmark was not to compete for the highest score, but to verify whether a from scratch proprietary model had reached a level at which it could compete on the global stage.

Having its name appear on the GAIA leaderboard is highly significant because it provides verified credibility granted by an independent third-party evaluation. This serves as objective evidence that can be used at every stage, from attracting investment and expanding overseas to B2B sales.

A CLEVI representative said, “A considerable number of the 63 official incorrect answers resulted from output format mismatches, errors in interpreting visual materials, and questions that could not be answered because information had been lost. These are issues of post-processing precision and the availability of external information, rather than fundamental limitations in logical reasoning. Because it is a proprietary model, CLEVI can improve it itself. This is precisely the structural advantage of a from scratch proprietary model. CLEVI’s achievement this time poses the question to Korea’s AI industry: ‘How long will we continue to rely on borrowed brains?’ CLEVI has chosen the more difficult path of developing from scratch and aims to prove the answer on the global stage,” the representative said.

Back to newsroom
CLEVI

Language and region

Machine-translated languages are marked. Availability follows the published site bundle.

136 languages

Recommended

1

East Asia

7

Southeast Asia

11

South Asia

18

Central Asia

5

Middle East and the Caucasus

10

Western and Southern Europe

16

Britain and Ireland

4

Northern Europe and the Baltics

10

Central Europe and the Balkans

14

Eastern Europe

5

East Africa and the Horn

8

West and Central Africa

9

Southern Africa

8

The Americas

5

The Pacific

5