News

CLEVI achieves 79.07% on GAIA with its own from scratch model... Top 2.5% among 3,090 models

Source: Electronic Times, Reporter Lee Won-ji

Source: Electronic Times, Reporter Lee Won-ji

Full quotation from the article “CLEVI achieves 79.07% on GAIA with its own from scratch model... Top 2.5% among 3,090 models”

Publication date: 2026-04-07

Link: https://www.etnews.com/20260407000203

cip-5.5-agent·cip-5.5-mm independent model stack: all five agents score 70 or above

98%+ accuracy when accounting for information loss, exceeding humans (92%)

Main image

AI start-up 'CLEVI' recorded an accuracy rate of 79.07% on the global AI agent benchmark 'GAIA' using its own from scratch model, placing it within the top 2.5% of all 3,090 registered models.

According to industry sources, GAIA is regarded in the AI benchmark market as accurately reflecting the capabilities of 'practical AI agents'. Co-developed by Meta AI, HuggingFace and Université Paris-Saclay, this benchmark was first released in November 2023.

GAIA is distinguished from other benchmarks in three key ways. First, overfitting is impossible because the answers are not disclosed. Second, high scores cannot be achieved through simple pattern matching because it requires multi-step, multimodal complex reasoning. Third, more than 3,090 models from around the world compete under identical conditions via the official Hugging Face leaderboard.

The test set consists of 301 questions in total. Level 1 involves using a single tool and simple reasoning; Level 2 involves combining multiple tools and intermediate-level reasoning; and Level 3 consists of the most difficult questions, requiring agent planning and execution across five or more steps. At the time the benchmark was announced, GPT models (with plugins) achieved only around 15%, while leading teams have now raised this to the 70–80% range.

CLEVI emphasised that the most technically noteworthy aspect of this achievement is that it did not use any external LLM APIs, such as GPT, Claude or Gemini. A key feature is that CLEVI used two models developed independently from scratch.

cip-5.5-agent is an agentic AI model that serves as the core brain of an agent pipeline that autonomously plans, executes and verifies complex tasks. cip-5.5-mm is a high-performance general-purpose multimodal model that understands and reasons over various file formats, including audio, images and video. As a significant proportion of GAIA questions include non-text inputs such as PDFs, image files and audio files, this multimodal capability became a key factor in achieving a high score.

Based on these two proprietary models, CLEVI entered GAIA with five different agent configurations, all of which scored above 70.

According to the company, an interesting result was identified during CLEVI’s internal post-review. For some of the 301 questions in total, the evidence supporting the correct answers had disappeared from the publicly available web. In these cases, information that had previously been available online or through search engines had been deleted or changed and was no longer accessible. When reassessed based on the information currently verifiable by people in the publicly available domain, CLEVI’s accuracy was found to be above 98%. This already exceeds the human average of 92%.

A CLEVI representative explained: “CLEVI achieved scores of 70+ with all five agents using only its own from scratch models and proprietary AI agent solutions, without mixing external models, placing it in the top 2.5% of the public benchmark. Further improvements to its proprietary models and agent solutions are expected to enable additional gains in the official score. This achievement is significant because it demonstrates ‘verified trust’ measured on a global public benchmark, represents an achievement by a domestic AI developer based on its own from scratch models, is the result of an independent model stack rather than a mixture of external models, and merits interpretation beyond the score itself, given the loss of information for some questions.”

Back to newsroom
CLEVI

Language and region

Machine-translated languages are marked. Availability follows the published site bundle.

136 languages

Recommended

1

East Asia

7

Southeast Asia

11

South Asia

18

Central Asia

5

Middle East and the Caucasus

10

Western and Southern Europe

16

Britain and Ireland

4

Northern Europe and the Baltics

10

Central Europe and the Balkans

14

Eastern Europe

5

East Africa and the Horn

8

West and Central Africa

9

Southern Africa

8

The Americas

5

The Pacific

5