Research

Clevi Semantic Database

One of the biggest concerns when AI is introduced into real-world work is the catastrophic forgetting phenomenon that occurs when previously learned knowledge is fine-tuned with new data.

1. Limitations of Catastrophic Forgetting and RAG

One of the biggest concerns when AI is introduced into real-world work is the catastrophic forgetting (Catastrophic Forgetting) phenomenon that occurs when previously learned knowledge is fine-tuned with new data.

This is a phenomenon in which neural network-based AI rapidly loses knowledge or patterns it previously learned while acquiring new information.

For example, fine-tuning an open-source LLM with data from a specific company may reduce its general knowledge or language capabilities.

To avoid this issue, RAG (Retrieval-Augmented Generation) technology is widely used, but RAG also has limitations.

Image in body

Key limitations of RAG

  • Insufficient structured data processing (Structured Data QA): RAG systems are primarily optimised for unstructured text documents, so they cannot properly understand or utilise the meaning and relationships of structured data (tables, databases, spreadsheets, etc.).
  • Limitations with complex PDFs/unstructured documents (Complex PDFs): Complex PDF documents (including tables, multiple columns and images) may lose information or have their context distorted during the chunking process. As a result, important data may not be indexed or may be difficult to retrieve.
  • Lack of a Fallback (backup) model (Fallback Model(s)): When a RAG system cannot find an appropriate answer, the logic for switching to an alternative (backup) model may be insufficient or inefficient. As a result, a satisfactory alternative may not be provided when answer generation fails.
  • Security vulnerabilities (LLM Security): Inadequate protection against security vulnerabilities in the LLM itself (such as prompt injection and data leakage) creates a risk of exposing sensitive information.
  • Insufficient scalability (Data Ingestion Scalability): Scalability issues arise during the indexing and storage of large volumes of data. As the amount of data increases, chunking, storage and retrieval speeds decrease, and the system load increases.
  • Missing important information (Missing Content): Important content is often omitted during the chunking process or not indexed. As a result, users cannot retrieve the information they want.
  • Missed top-ranked results (Missed Top Ranked): The chunk that is actually the most relevant (top-ranked) may be omitted from the search results. Causes include limitations of the search algorithm, reduced embedding quality and ranking errors.
  • Context mismatch (Not in Context): Retrieved chunks are often not aligned with the actual context of the question. This is a limitation of simple similarity-based retrieval, which lacks semantic connectedness.
  • Format errors (Wrong Format): The format of a retrieved chunk may be unsuitable for processing by the LLM or may differ from the answer format the user wants. For example, information may be distorted when a table is converted into text.
  • Information extraction failure (Not Extracted): Required information may not be properly extracted from a chunk, or the LLM may fail to recognise it. This often occurs with complex sentence structures, tables and images.
  • Incorrect information scope/depth (Incorrect Specificity): An answer may be too broad or, conversely, excessively specific for the question, meaning it does not align with the user's actual requirements. Adjusting the scope and depth of information is difficult.
  • Incomplete answer (Incomplete): The final generated answer may be incomplete or provide only some of the information, failing to sufficiently meet the user's requirements. Causes include an inability to combine information from multiple chunks or the LLM using only some of the available information.

As RAG relies too heavily on simple vector similarity searches and chunk-based indexing, its limitations in terms of reliability, scalability and accuracy are clear in real-world business applications.

2. CLEVI’s Semantic Database Overcoming the Limitations of RAG

To overcome these limitations, CLEVI developed the next-generation AI knowledge infrastructure, Clevi Semantic Database, optimised for semantic connections, complex reasoning and evidence-based responses.

This solution is possible because CLEVI possesses its own Reasoning models, multimodal AI and agentic AI models, and is one of CLEVI’s powerful technologies for making AI genuinely useful in the real world.

By analogy, if a database containing information is a library, CLEVI’s Semantic Database can be thought of as having an outstanding librarian. (When you ask the librarian for the information you need, you can receive an accurate and prompt response.)

Users can add new information to the library and retrieve the information they need at any time.

They can also configure data version management and access permissions as desired.

Every action performed by the librarian is recorded in logs, so it can be corrected even if a mistake occurs.

Main image

<Clevi Semantic Database Structure>

Key features and technical structure

Semantic data storage

  • Stores semantic relationships between data (context, hierarchy, causality, etc.) in a graph DB, rather than using a simple chunk vector DB
  • Easy to expand and supplement in the form of a knowledge graph/semantic network

Hybrid search and reasoning

  • CTA+Context Query: Reflects the query’s context and CTA together
  • Hybrid Retrieval: Combines vector similarity with rule-/graph-based searches
  • Intelligent Folder Hits: Automatically explores semantically related folders/categories

Simultaneous multimodal processing

  • Simultaneously interprets various types of data, including text, images, tables and audio, through TextReader Pool, VisionReader Pool, etc.
  • While existing RAG can primarily process text, the Semantic Database demonstrates outstanding multimodal processing capabilities

Evidence-based responses and reliability verification

  • Doc Summaries & Citations, automatic generation of document summaries and sources
  • Merge & Verify: semantic integration and reliability verification of information from multiple sources
  • Evidence Pack: provision of evidence-based response packages

Complex Reasoning and Planner

  • Planner/DAG: step-by-step planning and DAG (Directed Acyclic Graph)-based reasoning for complex queries

Integrated Agent Architecture

  • cip-5-agent Supervisor: top-level agent that manages and coordinates the entire search, reasoning and integration process
Images in the body

3. Structural Summary and Comparison

In summary, Semantic DB receives data in various forms, including audio, text, images and video, and classifies knowledge by connecting not only simple chunks but also semantic relationships between data, such as context, hierarchy and causality.

Based on this, it performs search and reasoning using semantic connections rather than keyword- or similarity-based searches such as RAG, enabling AI to provide more accurate information retrieval and answers.

In addition, because it is stored in the form of a knowledge graph or semantic network, it can be continuously expanded and supplemented with ease.

At Clevi, we identified the limitations of RAG systems in the era of AI transformation and, through a semantic database and proprietary AI models that address these limitations, we can now respond to customers quickly with more accurate and reliable data.

Clevi's AI system can create value from customers' valuable data in any environment, regardless of the type of domain.

Clevi will continue to do its utmost to maximise the value of our customers' data and provide innovative AI experiences.

We invite you to experience the new future of AI with Clevi.

Enquiries and demo requests: [email protected]

Copyright© 2025 Clevi Inc. All rights reserved.

Back to newsroom
CLEVI

Language and region

Machine-translated languages are marked. Availability follows the published site bundle.

136 languages

Recommended

1

East Asia

7

Southeast Asia

11

South Asia

18

Central Asia

5

Middle East and the Caucasus

10

Western and Southern Europe

16

Britain and Ireland

4

Northern Europe and the Baltics

10

Central Europe and the Balkans

14

Eastern Europe

5

East Africa and the Horn

8

West and Central Africa

9

Southern Africa

8

The Americas

5

The Pacific

5