News

Data-proven competitiveness—Domestic AI agent achieves top 2.5% on the GAIA benchmark

When the GAIA benchmark was first released, even GPT-4, which was the leading model at the time, scored only around 30%. This meant that, given GAIA’s requirements for complex reasoning, tool use, and multi-step problem-solving beyond simple question answering, AI at the time was still in its early stages when it came to “thinking and taking action”…

What GAIA scores mean — How far have AI agents advanced?

When the GAIA benchmark was first released, even GPT-4, which was the leading model at the time, scored only around 30%. This meant that, given GAIA’s requirements for complex reasoning, tool use, and multi-step problem-solving beyond simple question answering, AI at the time was still in its early stages when it came to the ability to “think and take action.”

However, leading agents have now reached 92.36%.

This change is more than a simple performance improvement. It is an indicator that AI has now moved beyond answering questions and entered the “Agentic AI” stage, in which it interprets situations, selects the tools it needs, and plans and executes multiple steps on its own.

In just a few years, AI has evolved from a “knowledge-generation tool” into an “agent that carries out work.”

Body image

📌 And CLEVI is among the top 2.5%

In this global competitive environment, the fact that CLEVI’s domestically developed agent was listed among the top 2.5% on the GAIA leaderboard demonstrates technological self-reliance and proves that an AI agent made in Korea can meet global standards. This means more than just a number. It signifies technological capability objectively validated through competition with thousands of models on a globally open benchmark, rather than in a specific demo environment.

More importantly, this achievement was not the accidental result of a single model. Agents with different designs repeatedly achieved top-tier performance on the same Agent Stack.

In other words, CLEVI’s technology is not a “one-time successful result,” but reproducible structural competitiveness and platform-level capability that can be extended across diverse industries and scenarios.

Body image

So, what does this indicator mean?

From the user’s perspective, the biggest barrier to adopting AI is the question, “Can I really trust it with my work?”

Performance repeatedly proven on the stage of a third-party global public leaderboard—not through flashy demos or company presentations—is the most objective answer to that question. Specifically, it is significant in three ways.

First, reduced adoption risk

Connecting unvalidated AI to internal operations represents a significant burden for enterprises.

Performance on credible benchmarks such as GAIA can be directly used as a basis for trust during the technology review stage.

Second, the realization of complex workflow automation

A top 2.5% ranking signifies a level of intelligence that goes beyond simple repetitive tasks to understand context, make decisions independently across multiple steps, and complete them.

Work areas previously considered "difficult to entrust to AI" can become practical targets for automation.

Third, achieving data security and performance simultaneously

The proprietary Agent Stack architecture means that data flows do not leave the organization.

This makes it possible to move beyond the dilemma of having to compromise security in order to use high-performance AI.

This architecture is a decisive differentiator, particularly in industries with highly sensitive data, such as finance, healthcare, and legal services.

The reproducibility of this architecture has already begun to be demonstrated.

And the performance of each agent built on that architecture will soon lead to tangible change in the field.

"Ultimately, the quality of technology is completed through 'confidence' in the field."

CLEVI does not simply provide high-performance AI. Our essence is to build an 'execution infrastructure' that breaks down the security barriers enterprises face and can actually complete complex business tasks.

Now, move beyond the stage of considering adoption and start a safe, powerful AI work environment with CLEVI.

As a reliable partner you can trust with your work, we will create meaningful innovation together in your business environment.

Back to newsroom
CLEVI

Language and region

Machine-translated languages are marked. Availability follows the published site bundle.

136 languages

Recommended

1

East Asia

7

Southeast Asia

11

South Asia

18

Central Asia

5

Middle East and the Caucasus

10

Western and Southern Europe

16

Britain and Ireland

4

Northern Europe and the Baltics

10

Central Europe and the Balkans

14

Eastern Europe

5

East Africa and the Horn

8

West and Central Africa

9

Southern Africa

8

The Americas

5

The Pacific

5
Data-proven competitiveness—Domestic AI agent achieves top 2.5% on the GAIA benchmark — CLEVI