Research

Clevi Semantic Database

One of the biggest concerns when AI is introduced into real-world operations is the phenomenon of catastrophic forgetting that occurs when fine-tuning previously learned knowledge with new data.

1. Catastrophic Forgetting and the Limitations of RAG

One of the biggest concerns when AI is introduced into real-world operations is the phenomenon of catastrophic forgetting (Catastrophic Forgetting) that occurs when fine-tuning previously learned knowledge with new data.

This is a phenomenon in which neural network-based AI rapidly loses previously acquired knowledge or patterns while learning new information.

For example, fine-tuning an open-source LLM with data from a specific organisation may result in reduced general knowledge or language capabilities.

To avoid this problem, RAG (Retrieval-Augmented Generation) technology is widely used, but RAG also has limitations.

Main text image

Key limitations of RAG

  • Insufficient structured data processing (Structured Data QA): RAG systems are primarily optimised for unstructured text documents, so they cannot properly understand or utilise the meaning and relationships of structured data (tables, databases, spreadsheets, etc.).
  • Limitations with complex PDFs/unstructured documents (Complex PDFs): Complex PDF documents (including tables, multiple columns and images) may lose information or have their context distorted during the chunking process. As a result, important data may not be indexed or may be difficult to search.
  • Lack of a Fallback model (Fallback Model(s)): When a RAG system cannot find an appropriate answer, the logic for switching to an alternative (backup) model may be insufficient or inefficient. As a result, a satisfactory alternative may not be provided to the user when answering fails.
  • Security vulnerabilities (LLM Security): Inadequate protection against security vulnerabilities in the LLM itself (such as prompt injection and data leakage) creates a risk of exposing sensitive information.
  • Insufficient scalability (Data Ingestion Scalability): Scalability issues arise during the indexing and storage of large volumes of data. As the amount of data increases, chunking, storage and search speeds decrease, while the system load increases.
  • Missing important information (Missing Content): Important content in documents is often omitted during the chunking process or not indexed. As a result, users cannot search for the information they need.
  • Missing top-ranked results (Missed Top Ranked): The chunk that is actually most relevant (the top-ranked chunk) may be omitted from the search results. Causes include limitations of the search algorithm, reduced embedding quality and ranking errors.
  • Context mismatch (Not in Context): The retrieved chunk is often unrelated to the context of the actual question. This is a limitation of simple similarity-based searches, which lack semantic connectivity.
  • Incorrect format (Wrong Format): The format of the retrieved chunk may be unsuitable for processing by the LLM or may differ from the answer format the user wants. For example, information may be distorted when a table is converted into text.
  • Information extraction failure (Not Extracted): The required information may not be properly extracted from the chunk, or the LLM may fail to recognise it. This often occurs with complex sentence structures, tables and images.
  • Inappropriate information scope/depth (Incorrect Specificity): The answer may be too broad for the question or, conversely, excessively specific, meaning it does not match the user's actual requirements. Adjusting the scope and depth of information is difficult.
  • Incomplete answer (Incomplete): The final answer generated may be incomplete or provide only some of the information, failing to fully satisfy the user's requirements. This may be caused by an inability to synthesise information from multiple chunks or by the LLM using only some of the information.

As RAG relies excessively on simple vector similarity searches and chunk-based indexing, its limitations in terms of reliability, scalability and accuracy are clear in real-world business environments.

2. Clevi’s Semantic Database Overcoming the Limitations of RAG

To overcome these limitations, CLEVI has developed the next-generation AI knowledge infrastructure, Clevi Semantic Database, optimised for semantic connections, complex reasoning and evidence-based responses.

This solution is possible because CLEVI possesses its own Reasoning models, multimodal AI and agentic AI models, and is one of CLEVI’s powerful technologies for enabling AI to provide real-world assistance.

By analogy, if a database containing information is considered a library, CLEVI’s Semantic Database can be thought of as having a highly capable librarian. (When you ask the librarian for the information you need, you can receive an accurate and prompt response.)

Users can add new information to the library and retrieve the information they need at any time.

They can also configure data version management and access permissions as desired.

As every action taken by the librarian is retained in logs (records), it can be corrected even if an error occurs.

Main image

<Clevi Semantic Database Structure>

Key Features and Technical Structure

Semantic data storage

  • Stores semantic relationships between data (context, hierarchy, causality, etc.) in a graph DB, rather than using a simple chunk vector DB
  • Easy to expand and supplement in the form of a knowledge graph/semantic network

Hybrid search and reasoning

  • CTA+Context Query: Reflects the query’s context and CTA together
  • Hybrid Retrieval: Combines vector similarity with rule-/graph-based searches
  • Intelligent Folder Hits: Automatically explores semantically related folders/categories

Simultaneous multimodal processing

  • Simultaneously interprets various types of data, including text, images, tables and audio, through TextReader Pool, VisionReader Pool and more
  • While conventional RAG can primarily process text, the Semantic Database excels in multimodal processing

Evidence-based responses and reliability verification

  • Doc Summaries & Citations, automatic generation of document summaries and sources
  • Merge & Verify: semantic integration and reliability verification of information from multiple sources
  • Evidence Pack: provision of evidence-based response packages

Complex Reasoning and Planner

  • Planner/DAG: step-by-step planning and DAG (Directed Acyclic Graph)-based reasoning for complex queries

Integrated Agent Architecture

  • cip-5-agent Supervisor: top-level agent that manages and coordinates the entire search, reasoning and integration process
Image in the main text

3. Structural Summary and Comparison

In summary, a Semantic DB takes in data in various forms (voice, text, images, video, etc.) and classifies knowledge by connecting not only simple chunks but also semantic relationships between data (context, hierarchy, causality, etc.).

Based on this, it performs search and reasoning using semantic connections rather than keyword- or similarity-based searches such as RAG, enabling AI to provide more accurate information retrieval and answers.

In addition, because it is stored in the form of a knowledge graph or semantic network, it is easy to continuously expand and supplement.

CLEVI identified the limitations of RAG systems in the era of AI transformation and, through a semantic database and proprietary AI models that can address them, enables customers to receive more accurate and reliable data quickly.

CLEVI’s AI systems can create value from customers’ valuable data in any environment, regardless of the type of domain.

CLEVI will continue to do its utmost to maximise the value of customers’ data and provide innovative AI experiences.

We invite you to experience the future of new AI with CLEVI.

Enquiries and demo requests: [email protected]

Copyright© 2025 Clevi Inc. All rights reserved.

Back to newsroom
CLEVI

Language and region

Machine-translated languages are marked. Availability follows the published site bundle.

136 languages

Recommended

1

East Asia

7

Southeast Asia

11

South Asia

18

Central Asia

5

Middle East and the Caucasus

10

Western and Southern Europe

16

Britain and Ireland

4

Northern Europe and the Baltics

10

Central Europe and the Balkans

14

Eastern Europe

5

East Africa and the Horn

8

West and Central Africa

9

Southern Africa

8

The Americas

5

The Pacific

5