1. Limitations of Catastrophic Forgetting and RAG
When AI is introduced into real-world operations, one of the biggest concerns is the phenomenon of Catastrophic Forgetting, which occurs when previously learned knowledge is fine-tuned with new data.
This is a phenomenon in which neural network-based AI rapidly loses previously learned knowledge or patterns while learning new information.
For example, fine-tuning an open-source LLM with data from a specific organisation may reduce its general knowledge or language capabilities.
To avoid this issue, RAG (Retrieval-Augmented Generation) technology is widely used, but RAG also has its limitations.
Key limitations of RAG
- Insufficient Structured Data Processing (Structured Data QA): RAG systems are primarily optimised for unstructured text documents, so they may fail to properly understand or utilise the meaning and relationships in structured data (tables, databases, spreadsheets, etc.).
- Limitations with Complex PDFs/Unstructured Documents (Complex PDFs): Complex PDF documents (including tables, multiple columns, images, etc.) may experience information loss or context distortion during the chunking process. As a result, important data may not be indexed or may become difficult to retrieve.
- Absence of a Fallback Model (Fallback Model(s)): When a RAG system cannot find an appropriate answer, the logic for switching to an alternative (fallback) model may be insufficient or inefficient. As a result, a satisfactory alternative may not be provided when answer generation fails.
- Security Vulnerabilities (LLM Security): Inadequate protection against security vulnerabilities in the LLM itself (prompt injection, data leakage, etc.) creates a risk of sensitive information being exposed.
- Insufficient Scalability (Data Ingestion Scalability): Scalability issues may arise during the indexing and storage of large volumes of data. As the amount of data increases, chunking, storage, and retrieval speed may decline, increasing the system load.
- Missing Important Information (Missing Content): Important content may often be omitted during the chunking process or may not be indexed. As a result, users may be unable to retrieve the information they need.
- Missing Top-Ranked Results (Missed Top Ranked): The most relevant chunk (top-ranked result) may be missing from the search results. Possible causes include limitations in the search algorithm, reduced embedding quality, and ranking errors.
- Context Mismatch (Not in Context): The retrieved chunk may often not match the context of the actual question. This is a limitation of simple similarity-based retrieval, which lacks semantic connectivity.
- Incorrect Format (Wrong Format): The format of the retrieved chunk may be unsuitable for processing by the LLM or may differ from the answer format expected by the user. For example, information may be distorted when a table is converted into text.
- Information Extraction Failure (Not Extracted): Required information may not be properly extracted from the chunk, or the LLM may fail to recognise it. This commonly occurs with complex sentence structures, tables, images, and similar content.
- Incorrect Information Scope/Depth (Incorrect Specificity): The answer may be too broad or, conversely, excessively specific for the question, and therefore may not match the user's actual requirements. Adjusting the scope and depth of information can be difficult.
- Incomplete Answer (Incomplete): The final generated answer may be incomplete or provide only partial information, failing to sufficiently meet the user's requirements. This may occur when information from multiple chunks is not consolidated or when the LLM uses only part of the available information.
As RAG relies excessively on simple vector similarity search and chunk-based indexing, its limitations in terms of reliability, scalability and accuracy are clear in actual business applications.
2. CLEVI’s Semantic Database that Overcomes the Limitations of RAG
To overcome these limitations, CLEVI has developed the next-generation AI knowledge infrastructure, Clevi Semantic Database, optimised for semantic connections, complex reasoning and evidence-based responses.
This solution is possible because CLEVI possesses its own Reasoning models, multimodal AI and agentic AI models, and is one of CLEVI’s powerful technologies that enables AI to provide real assistance in the real world.
By analogy, if the database where information is stored is a library, CLEVI’s Semantic Database can be thought of as having a highly capable librarian. (When you request the information you need from the librarian, you can receive an accurate and prompt response.)
Users can add new information to the library and retrieve the information they need at any time.
You can also configure data version management and access permissions as desired.
As every action taken by the librarian is retained as a log (record), it can be corrected even if an error occurs.
<Clevi Semantic Database Structure>
Key Features and Technical Structure
Semantic Data Storage
- Stores semantic relationships between data (context, hierarchy, causality, etc.) in a graph DB, rather than using a simple chunk vector DB
- Easy to expand and enhance in the form of a knowledge graph/semantic network
Hybrid Search and Reasoning
- CTA+Context Query: Reflects the context of the query and the CTA together
- Hybrid Retrieval: Combines vector similarity and rule/graph-based search
- Intelligent Folder Hits: Automatically explores semantically related folders/categories
Simultaneous Multimodal Processing
- Simultaneously interprets diverse data, including text, images, tables and audio, using TextReader Pool, VisionReader Pool and other tools
- While existing RAG can mainly process text, the Semantic Database excels at multimodal processing
Evidence-Based Responses and Reliability Verification
- Doc Summaries & Citations, automatic generation of document summaries and sources
- Merge & Verify: Semantic integration and reliability verification of information from multiple sources
- Evidence Pack: Providing evidence-based response packages
Complex Reasoning and Planner
- Planner/DAG: Step-by-step planning and DAG (Directed Acyclic Graph)-based reasoning for complex queries
Integrated Agent Architecture
- cip-5-agent Supervisor: The top-level agent that manages and coordinates the entire search, reasoning and integration process
3. Structural Summary and Comparison
In summary, Semantic DB accepts data in various forms, such as voice, text, images and video, and classifies knowledge by connecting not only simple chunks but also the semantic relationships between data, including context, hierarchy and causality.
Based on this, it enables search and reasoning using semantic connections rather than keyword- or similarity-based searches such as RAG, allowing AI to provide more accurate information retrieval and answers.
Moreover, as the data is stored in the form of a knowledge graph or semantic network, it can be continuously expanded and enhanced with ease.
At Clevi, we identified the limitations of RAG systems in the era of AI transformation and, through a semantic database and proprietary AI models that address these limitations, we can now provide customers with faster responses based on more accurate and reliable data.
Clevi’s AI system can make customers’ valuable data useful, regardless of the domain, in any environment.
Clevi will continue to do its utmost to maximise the value of customers’ data and provide innovative AI experiences.
We invite you to experience the future of a new era of AI with Clevi.
Enquiries and Demo Requests: [email protected]
Copyright© 2025 Clevi Inc. All rights reserved.
