Research

Clevi Semantic Database

One of the biggest concerns when AI is introduced into real-world business operations is the phenomenon of catastrophic forgetting that occurs when previously learned knowledge is fine-tuned with new data.

1. Catastrophic Forgetting and the Limitations of RAG

One of the biggest concerns when AI is introduced into real-world business operations is the phenomenon of catastrophic forgetting (Catastrophic Forgetting) that occurs when previously learned knowledge is fine-tuned with new data.

This is a phenomenon in which neural network-based AI rapidly loses previously learned knowledge or patterns while learning new information.

For example, fine-tuning an open-source LLM with data from a specific company may reduce its general knowledge or language abilities.

To avoid this issue, RAG (Retrieval-Augmented Generation) technology is widely used, but RAG also has limitations.

Body image

Representative limitations of RAG

  • Insufficient Structured Data Processing (Structured Data QA): RAG systems are primarily optimized for unstructured text documents, so they may fail to properly understand or utilize the meaning and relationships of structured data (tables, databases, spreadsheets, etc.).
  • Limitations with Complex PDFs/Unstructured Documents (Complex PDFs): Complex PDF documents (including tables, multiple columns, images, etc.) may lose information or have their context distorted during the chunking (segmentation) process. As a result, important data may not be indexed or may be difficult to retrieve.
  • Absence of a Fallback Model (Fallback Model(s)): When a RAG system cannot find an appropriate answer, its logic for switching to an alternative (fallback) model may be insufficient or inefficient. As a result, it may fail to provide a satisfactory alternative when an answer cannot be generated.
  • Security Vulnerabilities (LLM Security): Inadequate protection against security vulnerabilities in the LLM itself (such as prompt injection and data leakage) creates a risk of exposing sensitive information.
  • Insufficient Scalability (Data Ingestion Scalability): Scalability issues can arise during the indexing and storage of large volumes of data. As the amount of data increases, chunking, storage, and retrieval speeds decline, increasing the system load.
  • Missing Important Information (Missing Content): Important content in documents is often omitted during the chunking process or not indexed. As a result, users may be unable to retrieve the information they need.
  • Missing Top-Ranked Results (Missed Top Ranked): The chunk with the highest relevance (top-ranked chunk) may be missing from the search results. Causes include limitations in the search algorithm, reduced embedding quality, and ranking errors.
  • Context Mismatch (Not in Context): Retrieved chunks often do not match the context of the actual question. This is a limitation of simple similarity-based retrieval, which lacks semantic connectivity.
  • Incorrect Format (Wrong Format): The format of a retrieved chunk may be unsuitable for processing by the LLM or may differ from the response format the user wants. For example, a table may be converted into text, distorting the information.
  • Information Extraction Failure (Not Extracted): Required information may not be properly extracted from a chunk, or the LLM may fail to recognize it. This frequently occurs with complex sentence structures, tables, and images.
  • Inappropriate Information Scope/Depth (Incorrect Specificity): An answer may be too broad or, conversely, overly specific for the question, making it unsuitable for the user's actual needs. Adjusting the scope and depth of information is difficult.
  • Incomplete Answer (Incomplete): The final generated answer may be incomplete or provide only some of the information, failing to fully meet the user's needs. Causes include an inability to synthesize information from multiple chunks or the LLM using only part of the available information.

Because RAG relies too heavily on simple vector similarity search and chunk-based indexing, its limitations in terms of reliability, scalability, and accuracy are clear in real-world work.

2. CLEVI’s Semantic Database That Overcomes RAG’s Limitations

To overcome these limitations, CLEVI developed the next-generation AI knowledge infrastructure, optimized for semantic connections, complex reasoning, and evidence-based responsesClevi Semantic Database.

This solution is possible because CLEVI possesses its own Reasoning models, multimodal AI, and agentic AI models, and it is one of CLEVI’s powerful technologies that makes AI genuinely useful in the real world.

By analogy, if a database containing information is a library, CLEVI’s semantic database can be thought of as having an excellent librarian. (When you ask the librarian for the information you need, you can receive an accurate and prompt response.)

Users can add new information to the library and retrieve the information they need at any time.

They can also configure data version management and access permissions as desired.

Because every action taken by the librarian is retained as a log (record), it can be corrected even if an error occurs.

Main image

<Clevi Semantic Database Structure>

Key Features and Technical Structure

Semantic Data Storage

  • Stores semantic relationships between data (context, hierarchy, causality, etc.) in a graph database rather than a simple chunk vector DB
  • Easy to expand and enhance in the form of a knowledge graph/semantic network

Hybrid Search and Reasoning

  • CTA+Context Query: Reflects the query’s context and CTA together
  • Hybrid Retrieval: Combines vector similarity with rule-/graph-based search
  • Intelligent Folder Hits: Automatically explores semantically related folders/categories

Simultaneous Multimodal Processing

  • Simultaneously interprets various types of data, including text, images, tables, and audio, through TextReader Pool, VisionReader Pool, and more
  • While conventional RAG can primarily process text, semantic databases excel at multimodal processing

Evidence-Based Responses and Reliability Verification

  • Doc Summaries & Citations, automatically generating document summaries and sources
  • Merge & Verify: semantic integration and reliability verification of information from multiple sources
  • Evidence Pack: providing evidence-based response packages

Complex Reasoning & Planner

  • Planner/DAG: step-by-step planning and DAG (Directed Acyclic Graph)-based reasoning for complex queries

Integrated Agent Architecture

  • cip-5-agent Supervisor: the top-level agent that manages and coordinates the entire search, reasoning, and integration process
Images in the body

3. Structural Summary and Comparison

In summary, Semantic DB accepts data in various forms, including voice, text, images, and video, and classifies knowledge by connecting not only simple chunks but also semantic relationships between data, such as context, hierarchy, and causality.

Based on this, it performs search and reasoning using semantic connections rather than keyword- or similarity-based searches like RAG, enabling AI to provide more accurate information retrieval and answers.

In addition, because it is stored in the form of a knowledge graph or semantic network, it is easy to continuously expand and enhance.

At CLEVI, we identified the limitations of RAG systems in the era of AI transformation and, through our semantic database and proprietary AI models that address these limitations, we can now provide customers with faster responses based on more accurate and reliable data.

CLEVI’s AI systems can make customers’ valuable data useful in any environment, regardless of the type of domain.

CLEVI will continue to do its utmost to maximize the value of customers’ data and deliver innovative AI experiences.

We invite you to experience the future of new AI with CLEVI.

Inquiries and Demo Requests: [email protected]

Copyright© 2025 Clevi Inc. All rights reserved.

Back to newsroom
CLEVI

Language and region

Machine-translated languages are marked. Availability follows the published site bundle.

136 languages

Recommended

1

East Asia

7

Southeast Asia

11

South Asia

18

Central Asia

5

Middle East and the Caucasus

10

Western and Southern Europe

16

Britain and Ireland

4

Northern Europe and the Baltics

10

Central Europe and the Balkans

14

Eastern Europe

5

East Africa and the Horn

8

West and Central Africa

9

Southern Africa

8

The Americas

5

The Pacific

5
Clevi Semantic Database — CLEVI