News

Competitiveness proven by data—domestic AI agent achieves a place in the top 2.5% of the GAIA benchmark

When the GAIA benchmark was first released, even GPT-4, which was then considered a state-of-the-art model, managed to score only around 30%. Given that GAIA requires complex reasoning, tool use and multi-step problem-solving beyond simple question answering, this meant that AI was still at an early stage in its ability to “think and act” at the time …

What GAIA scores mean — How far have AI agents advanced?

When the GAIA benchmark was first released, even GPT-4, which was then considered a state-of-the-art model, managed to score only around 30%. Given that GAIA requires complex reasoning, tool use and multi-step problem-solving beyond simple question answering, this meant that AI was still at an early stage in its ability to “think and act” at the time.

However, today, top-performing agents have reached 92.36%.

This change is not simply an improvement in performance. It is an indicator that AI has now moved beyond answering questions and entered the ‘Agentic AI’ stage, where it interprets situations, selects the tools it needs, and independently plans and executes multiple steps.

In just a few years, AI has evolved from a “knowledge-generation tool” into an “agent that carries out work”.

Body image

📌 And CLEVI is among the top 2.5%

In this global competitive landscape, the fact that CLEVI’s agent, developed in Korea, has been listed among the top 2.5% on the GAIA leaderboard demonstrates technological self-reliance and proves that an AI agent made in Korea can meet global standards. This means more than just a number. It signifies technological capability objectively validated through competition with thousands of models on a publicly available global benchmark, rather than in a specific demo environment.

More importantly, this achievement was not the accidental result of a single model; agents with different designs repeatedly achieved top-tier performance on the same Agent Stack.

In other words, CLEVI’s technology is not a “one-off successful result” but a reproducible structural competitive advantage, with platform-level capabilities that can be expanded across various industries and scenarios.

Body image

So, what does this metric mean?

From the user’s perspective, the biggest barrier to adopting AI is the question: "Can I really trust it and entrust it with my work?"

Performance repeatedly proven on the stage of a global public leaderboard, an independent third party, rather than through flashy demos or company presentations, is the most objective answer to that question. Specifically, it is significant in three respects.

First, reducing adoption risk

Connecting unvalidated AI to internal operations is a considerable burden for enterprises.

Performance on credible benchmarks such as GAIA can be directly used as a basis for trust during the technology review stage.

Second, making complex work automation a reality

A ranking in the top 2.5% signifies a level of intelligence that goes beyond simple repetitive tasks, understanding context, making decisions independently across multiple steps, and completing them.

Areas of work that were previously considered "difficult to entrust to AI" can become practical targets for automation.

Third, securing data protection and performance simultaneously

The proprietary Agent Stack architecture means that data flows do not leave the organisation.

This allows enterprises to move beyond the existing dilemma of having to compromise security in order to use high-performance AI.

This architecture is a decisive differentiator, particularly in industries with highly sensitive data, such as finance, healthcare, and legal services.

The reproducibility of this architecture has already begun to be demonstrated.

And the performance of each agent built on this architecture will soon lead to tangible changes in the field.

"Ultimately, technological excellence is completed through 'confidence' in the field."

CLEVI does not simply provide high-performance AI. Our essence lies in building an 'execution-oriented infrastructure' that breaks down the security barriers enterprises face and can actually complete complex business tasks.

Now, move beyond the stage of considering adoption and start a safe, powerful AI work environment with CLEVI.

As a reliable partner you can confidently entrust your work to, we will work with you to create tangible innovation in your business operations.

Back to newsroom
CLEVI

Language and region

Machine-translated languages are marked. Availability follows the published site bundle.

136 languages

Recommended

1

East Asia

7

Southeast Asia

11

South Asia

18

Central Asia

5

Middle East and the Caucasus

10

Western and Southern Europe

16

Britain and Ireland

4

Northern Europe and the Baltics

10

Central Europe and the Balkans

14

Eastern Europe

5

East Africa and the Horn

8

West and Central Africa

9

Southern Africa

8

The Americas

5

The Pacific

5