News

AI startup CLEVI enters the top 2.5% of GAIA… demonstrating proven credibility

Source: Digital Times, Reporter Bon-gyu Koo

Source: Digital Times, Reporter Bon-gyu Koo

Full article citation: “AI startup CLEVI enters the top 2.5% of GAIA… demonstrating proven credibility”

Publication date: 2026-04-08

Link: https://n.news.naver.com/article/029/0003020511?sid=105

Korean AI startup ‘CLEVI’ announced that, while building its own model and agent solution developed from scratch, it became the first in Korea to enter the top 2.5% of the GAIA benchmark, based on all 3,090 registered models.

According to industry sources, GAIA, designed by Meta AI and operated by Hugging Face, evaluates not an AI’s ability to ‘do just one thing well,’ but its ability to ‘combine multiple capabilities to solve real-world problems.’ It must find information on the web, read tables in PDFs, analyse images, run code to perform calculations, and synthesize the results into a single answer. With 301 private questions, three difficulty levels, and a policy of keeping the answers confidential, it is structured to make both overfitting and cheating impossible.

Two things are required to achieve a high score on this test. The first is the reasoning ability of the underlying language model, and the second is an agent solution that connects the model to real-world tools and enables it to perform tasks autonomously. No matter how good the model is, a weak agent cannot solve Level 3; and no matter how sophisticated the agent is, insufficient reasoning ability causes incorrect answers to propagate at the intermediate stages.

The current structure of the GAIA leaderboard reveals an interesting pattern. Most top-ranked agents have adopted a ‘multi-model mix’ strategy that combines multiple Big Tech models, such as GPT, Claude, and Gemini. It is a rational approach that combines each model’s strengths and helps guard against score reductions caused by information loss and other factors.

CLEVI, however, has not joined that group. Developed from scratch, cip-5.5-agent (agentic AI) autonomously plans, executes, and verifies complex tasks, while cip-5.5-mm (high-performance general-purpose multimodal) understands and reasons over various file formats, including audio, images, and video. Both models were independently developed within CLEVI, and no external LLM API was used whatsoever.

With this proprietary model stack, CLEVI entered GAIA with five agents, all recording scores in the 70s or higher. From a top score of 79.07% to 70.76%, this demonstrated the stability of the entire stack—not the luck of a single result.

Image in the main text

According to the company, another figure lies behind the official score of 79.07%. A post-evaluation conducted internally by CLEVI found that, among the 301 questions, a considerable number no longer had supporting evidence for the correct answers available on the public web. When reevaluated based only on questions for which the correct answer still exists online, CLEVI’s accuracy rate was over 98%. This figure already exceeds the human average of 92%.

Of course, mixing in powerful general-purpose models from other companies could have raised the official score further. However, CLEVI deliberately did not choose that strategy. This is because the goal of this benchmark was not to compete for the highest score, but to verify whether a from scratch proprietary model had reached a level capable of competing on the global stage.

Having its name listed on the GAIA leaderboard is highly significant because it provides verified credibility conferred through an independent third-party evaluation. This serves as objective evidence that can be leveraged in every area, from attracting investment and expanding overseas to B2B sales.

A CLEVI representative said, “A considerable number of the 63 official incorrect answers resulted from output format mismatches, errors in interpreting visual materials, and questions that could not be answered because information had been lost. These are issues of post-processing precision and the availability of external information, rather than fundamental limitations in logical reasoning. Because it is a proprietary model, CLEVI can improve it independently. This is precisely the structural advantage of a from scratch proprietary model. CLEVI’s achievement this time poses the following question to the Korean AI industry: ‘How long will we continue to depend on borrowed brains?’ CLEVI has chosen the more difficult path of developing a from scratch model and aims to prove the answer on the global stage,” the representative stated.

Back to newsroom
CLEVI

Language and region

Machine-translated languages are marked. Availability follows the published site bundle.

136 languages

Recommended

1

East Asia

7

Southeast Asia

11

South Asia

18

Central Asia

5

Middle East and the Caucasus

10

Western and Southern Europe

16

Britain and Ireland

4

Northern Europe and the Baltics

10

Central Europe and the Balkans

14

Eastern Europe

5

East Africa and the Horn

8

West and Central Africa

9

Southern Africa

8

The Americas

5

The Pacific

5