“How could a small company create a world-class large language model?”
How did a small start-up create a large language model?
This has been the question CLEVI has received most frequently since announcing its 1.4 trillion-parameter Foundation Model.
Unlike many companies that simply fine-tune open-source models, CLEVI independently carried out the entire process, from data collection and model design to large-scale distributed training and quality validation.
Below, we introduce these secrets and competitive advantages in detail.
1. A small company creates a large language model?
When CLEVI released its self-developed large language model to the market in July 2025, numerous customers (users) were abuzz with surprise and scepticism. “Is it really self-developed?” “Did they just slightly modify an open-source model?”
In fact, many companies in Korea and overseas simply perform transfer learning (fine-tuning) on open-source models and promote them as their own models.
However, CLEVI positioned it as a **‘From-scratch Foundation Model’**.
In other words, it designed and executed the entire process itself, from data collection and model design to large-scale distributed training and quality validation.
2. Why developing a large language model is difficult
To train a From-Scratch Foundation Model, all of the following must be learned.
A large-scale, high-quality dataset is required
- Diversity: Securing diversity in languages, domains and formats (text, images, etc.)
- Curation: Removing noise, managing quality, and filtering duplicate/harmful data
- Licensing: Clarifying copyright and data usage rights
Large-scale computing infrastructure
- High-performance GPU/TPU clusters: Technology to integrate equipment across thousands to tens of thousands of nodes to support large-scale distributed training
- Efficient distributed training frameworks: DeepSpeed, Megatron-LM, Ray, etc.
- Storage/network: High-volume data input and output, and a high-speed network environment
Model architecture design
- Reflect the latest architecture: Transformer, MoE (Mixture of Experts), multimodal capabilities, and more
- Scalability: Consider scalability in terms of the number of parameters, layers, input length, and more
- Efficiency: Optimise training/inference speed and memory usage
Training strategies and algorithms
- Pre-training objectives: Language models (e.g. next token prediction), multimodal models (e.g. contrastive learning), and more
- Optimisation techniques: Mixed precision, gradient accumulation, learning rate schedules, and more
- Normalisation/stabilisation: Dropout, layer norm, weight decay, and more
Evaluation and validation framework
- Benchmark sets: Use standard datasets (Benchmarks)
- Internal evaluation: Evaluate based on real-world usage scenarios
- Continuous monitoring: Check quality, bias, and safety during and after training
Workforce and organisational capabilities
- AI/ML researchers: Model design and algorithm development
- Data engineers: Data collection, cleansing, and management
- Infrastructure engineers: Distributed systems and cloud/on-premises management
- Project managers: Schedule, budget, and quality management
Ethical, legal, and security considerations
- AI ethics: Bias, safety, transparency, and accountability
- Legal compliance: Privacy, copyright, and data sovereignty
- Security: Preventing data/model leakage and controlling access
There are currently around 2,000 AI researchers worldwide, most of whom are concentrated in big tech companies such as META, Google, and XAI. In South Korea, teams capable of directly designing everything from data collection to large-scale training infrastructure and creating a Foundation model are extremely rare.
3. Building a large language model with a small team.
However, Hwan-ho Lee, CEO of CLEVI, decided to create the world’s best AI at “CLEVI”.
Based on more than 10 years of experience researching deep learning and machine learning, CEO Hwan-ho Lee founded the company (CLEVI) three years ago with the goal of “creating the world’s best AI”.
By directly designing every process—including the company’s own data pipeline, large-scale distributed deep learning infrastructure, model architecture development, and quality validation—the company developed the capabilities required to build a Foundation Model.
After training models several times, the company now has the Clevi-5-x model, which can rival the world’s leading models.
To effectively train a large language model, infrastructure is required that connects hundreds to thousands of GPUs into a cluster and can process tens of quadrillions of computations simultaneously per second. In this process, technology that minimises computation synchronisation and communication latency in large-scale distributed environments is critical. Until now, most companies in South Korea have been unable to train models properly due to the lack of “precision control algorithms”. By independently developing this proprietary distributed computation control algorithm, Clevi CEO Lee Hwan-ho successfully trained South Korea’s first large language model with 1.4 trillion (1.4 Trillion) parameters.
4. Why a small number of developers could develop various models
Most big tech companies also have large-scale infrastructure control and data refinement technologies.
Hundreds to thousands of research personnel are required to manage this.
However, Clevi has developed and maintains various models with a small number of research personnel.
How was this possible?
Clevi develops and maintains various models with a small number of personnel through Teacher-Student Model(Knowledge Distilation) technology.
Let us take a look at Teacher-Student technology.
Teacher-Student (Knowledge Distillation)
Teacher-Student technology is, simply put, a technology that rapidly develops various models with a small number of personnel by evaluating and providing feedback on a child model (Student Model) through a large language model (Teacher Model), replacing hundreds of research personnel.
- The Mother Model (large-scale model) evaluates and provides feedback on models in various fields, replacing hundreds of research personnel.
- If training is conducted through Clevi-x-platform in this way, high-performance models such as Reasoning, VLM, Physical AI and Coding Agent can be trained and deployed in 10–15 days.
Clevi Coding Agent
Clevi Phsycal AI Learning Platform
VLM-Vision Language Model
Security Agent
5. Differentiation in Clevi-x-platform technology and quality
Model performance and scalability
- Reasoning (CoT/Reasoning) advantage: cip-5-x incorporates step-by-step logical reasoning and the presentation of supporting evidence into its design philosophy, demonstrating highly reliable computation and interpretation capabilities in solving scientific, mathematical and engineering problems. Based on internal benchmarks, it is being designed and operated with the goal of achieving performance at or above the GPT-5 level, while reproducibility and verifiability under identical prompt conditions are managed as core quality indicators.
- Full-spectrum multimodality: By producing and combining a purpose-driven VLM (cip-5-vision), a large-scale multimodal processing model (ivy-4-mm) and a lightweight on-device model (ivy-3-text) on a single platform, it enables optimal combinations tailored to task characteristics (search/classification/summarisation vs. reasoning/explanation/execution).
Enterprise applicability
- On-premises security and availability: Data and model parameters are separately protected through operation across the entire closed network (input–training–service–backup), HA design, modular scalability and SVCE (isolated secure execution environment).
- Rapid adoption and scaling: Ivy Chat (business-oriented multimodal conversations/agents) and API Platform (globally standardised SaaS) support rapid implementation and operation.
Physical AI differentiation
- By combining NVIDIA Isaac-based simulation with Evolutionary RL, and using Sim2Real transfer and parallel synthetic-data training, learning efficiency and field adaptability have been enhanced. Few alternatives cover the entire process from digital and robotics learning through to deployment.
Service quality and governance
- Model/prompt/tool version control, evaluation suites (Korean and English knowledge, industry-specific tasks, hallucination and safety indicators), change management (SLO/SLA) and operational observability (latency, success rate and TCO) are managed through a single-platform process.
There are currently more than approximately 300 AI model service companies in South Korea, but
CLEVI is the only company that can provide both on-premises and cloud services through high-performance, proprietary foundation models.
6. How to utilise language models in industry
As adopting AI requires substantial investment and rigorous validation, it is essential to directly compare and experience solutions to determine whether they will genuinely benefit your business.
- CLEVI goes beyond simply building models to provide AI solutions that can be applied in real-world industries.
- For example, it has a complete range of enterprise offerings, including Ivy Chat Platform, Clevi API Platform, industry-specific customisation, and support for security requirements.
- CLEVI has technical capabilities that are difficult to find domestically or internationally, including physical AI, robotics, and multimodal data processing.
CLEVI provides tangible AI solutions that use AI as a ‘means’ rather than an ‘end’, contributing to real-world work efficiency and industry innovation. Experience the distinctiveness of a genuine From-scratch Foundation model for yourself.
Enquiries and demo requests: [email protected]
Copyright© 2025 Clevi Inc. All rights reserved.
