News

AI startup Clevi enters the top 2.5% of GAIA… demonstrating verified credibility

Source: Digital Times, Reporter Gubon-gyu

Source: Digital Times, Reporter Gubon-gyu

Full article citation: “AI startup Clevi enters the top 2.5% of GAIA… demonstrating verified credibility”

Publication date: 2026-04-08

Link: https://n.news.naver.com/article/029/0003020511?sid=105

Clevi, a domestic AI startup, announced that after building its own model and proprietary agent solution from scratch, it became the first in South Korea to enter the top 2.5% of the GAIA benchmark, based on all 3,090 registered models.

According to industry sources, GAIA, designed by Meta AI and operated by Hugging Face, evaluates not an AI’s ability to ‘do just one thing well’, but its ability to ‘combine multiple capabilities to solve real-world problems’. It must find information on the web, read tables in PDFs, analyse images, run code to perform calculations, and combine the results to produce a single correct answer. With 301 private questions, three levels of difficulty, and a policy of keeping answers confidential, it is structured to make both overfitting and cheating impossible.

Two things are required to achieve a high score on this test. The first is the reasoning ability of the underlying language model, and the second is an agent solution that connects the model to real-world tools and enables it to perform tasks autonomously. No matter how good the model is, a weak agent cannot solve Level 3; and no matter how sophisticated the agent is, insufficient reasoning ability in the model causes incorrect answers to propagate from the intermediate stages.

The current structure of the GAIA leaderboard reveals an interesting pattern. Most of the top-ranked agents have adopted a ‘multi-model mix’ strategy that combines several big-tech models, including GPT, Claude and Gemini. This is a rational approach that combines the strengths of each model and safeguards against score reductions caused by information loss and other factors.

Clevi, however, has not joined that group. Developed from scratch, cip-5.5-agent (agentic AI) autonomously plans, executes and verifies complex tasks, while cip-5.5-mm (high-performance general-purpose multimodal) understands and reasons across various file formats, including audio, images and video. Both models were independently developed internally by Clevi, and no external LLM APIs were used at all.

And with this proprietary model stack, it entered five agents in GAIA, all recording scores of 70 or above. From the highest score of 79.07% to 70.76%, this demonstrated the stability of the entire stack, rather than the luck of a single result.

Body image

According to the company, there is another figure behind the official score of 79.07%. CLEVI’s internal post-evaluation found that, among the 301 questions, a considerable number had lost their sources of evidence for the correct answers on the publicly available web. When re-evaluated based only on questions whose answers still exist on the web, CLEVI’s accuracy rate was above 98%. This figure already exceeds the human average of 92%.

Of course, mixing in powerful general-purpose models from other companies could have raised the official score further. However, CLEVI deliberately did not choose that strategy. This was because the goal of this benchmark was not to compete for the highest score, but to verify whether a from scratch proprietary model had reached a level at which it could compete on the global stage.

Having its name listed on the GAIA leaderboard is highly significant because it demonstrates verified credibility awarded through an independent third-party evaluation. This provides objective evidence that can be used at every stage, from attracting investment and expanding overseas to B2B sales.

A CLEVI representative said, “Of the 63 officially incorrect answers, many resulted from output format mismatches, errors in interpreting visual materials, and questions that could not be answered because information had been lost. These are issues of post-processing precision and the availability of external information, rather than fundamental limitations in logical reasoning. Because it is a proprietary model, CLEVI can improve it independently. This is precisely the structural advantage of a from scratch proprietary model. CLEVI’s achievement this time poses the question to Korea’s AI industry: ‘How long will we continue to depend on borrowed brains?’ CLEVI has chosen the more difficult path of developing a from scratch model and aims to prove the answer on the global stage,” the representative said.

Back to newsroom
CLEVI

Language and region

Machine-translated languages are marked. Availability follows the published site bundle.

136 languages

Recommended

1

East Asia

7

Southeast Asia

11

South Asia

18

Central Asia

5

Middle East and the Caucasus

10

Western and Southern Europe

16

Britain and Ireland

4

Northern Europe and the Baltics

10

Central Europe and the Balkans

14

Eastern Europe

5

East Africa and the Horn

8

West and Central Africa

9

Southern Africa

8

The Americas

5

The Pacific

5
AI startup Clevi enters the top 2.5% of GAIA… demonstrating verified credibility — CLEVI