Ensiklopedia VibeKoding: AI Capability Dictionary.Ensiklopedia VibeKoding: AI Capability Dictionary.
As generative AI technologies become widely adopted across various products and business scenarios, an increasingly practical question confronts each of us: What AI capabilities are actually available? And for a specific requirement, which capability, which type of model, or which product should be chosen to implement it?As generative AI technologies become widely adopted across various products and business scenarios, an increasingly practical question confronts each of us: What AI capabilities are actually available? And for a specific requirement, which capability, which type of model, or which product should be chosen to implement it?
Faced with this confusion, the most intuitive approach might be "cramming at the last minute": search for cloud service providers' product APIs or corresponding models when a need arises, then look up commercial solutions, compare documentation and demos, and proceed. When you see an image-related requirement, you think of image generation; when you encounter a text task, you pull in a large model; when it involves voice interaction, you recall ASR and TTS — and then shop around among a sea of APIs and services. However, cobbling together scattered products is fundamentally different from systematically planning, selecting, and combining AI capabilities in enterprise-level scenarios. Relying solely on ad-hoc research and experience-based judgment leads to a series of serious challenges: fragmented capability awareness, arbitrary solution design, and difficulty in reusing capabilities.Faced with this confusion, the most intuitive approach might be "cramming at the last minute": search for cloud service providers' product APIs or corresponding models when a need arises, then look up commercial solutions, compare documentation and demos, and proceed. When you see an image-related requirement, you think of image generation; when you encounter a text task, you pull in a large model; when it involves voice interaction, you recall ASR and TTS — and then shop around among a sea of APIs and services. However, cobbling together scattered products is fundamentally different from systematically planning, selecting, and combining AI capabilities in enterprise-level scenarios. Relying solely on ad-hoc research and experience-based judgment leads to a series of serious challenges: fragmented capability awareness, arbitrary solution design, and difficulty in reusing capabilities.
To address these pain points, this article is organized around the core idea of an "AI Capability Panorama." In this handbook, our goal is not to pile up jargon, but to help you quickly figure out three things: "What AI capability can handle this task? Which type of model or product should I roughly choose? What keywords should I use next to find APIs, projects, or services to try out?" Through a systematic review spanning modalities (text, image, audio, video, 3D, multimodal) and architecture layers (models, retrieval, agents, platform engineering), we can identify the corresponding AI capabilities, representative models/products, and common real-world business use cases for each typical requirement and scenario, helping teams build AI systems with lower trial-and-error costs, higher decision-making efficiency, and stronger reusability.To address these pain points, this article is organized around the core idea of an "AI Capability Panorama." In this handbook, our goal is not to pile up jargon, but to help you quickly figure out three things: "What AI capability can handle this task? Which type of model or product should I roughly choose? What keywords should I use next to find APIs, projects, or services to try out?" Through a systematic review spanning modalities (text, image, audio, video, 3D, multimodal) and architecture layers (models, retrieval, agents, platform engineering), we can identify the corresponding AI capabilities, representative models/products, and common real-world business use cases for each typical requirement and scenario, helping teams build AI systems with lower trial-and-error costs, higher decision-making efficiency, and stronger reusability.
In this handbook, we will systematically introduce the current mainstream AI capability landscape — from single modalities to multimodal fusion, from individual models to the overall framework of platforms and engineering — combined with common product forms and application scenarios, to provide practice-oriented capability selection references.In this handbook, we will systematically introduce the current mainstream AI capability landscape — from single modalities to multimodal fusion, from individual models to the overall framework of platforms and engineering — combined with common product forms and application scenarios, to provide practice-oriented capability selection references.
> Due to the extensive content, you may consult the handbook when you encounter scenarios in practice where you're unsure how to select capabilities. It is recommended that you let AI reference this handbook based on specific application directions and provide suggested model selection recommendations and solution API calling advice.> Due to the extensive content, you may consult the handbook when you encounter scenarios in practice where you're unsure how to select capabilities. It is recommended that you let AI reference this handbook based on specific application directions and provide suggested model selection recommendations and solution API calling advice.
If you only want to understand the corresponding categories without reading the detailed content, just read the opening paragraph of each major section, such as 1.1 and 1.2, but you don't need to read 1.1.1 or 1.1.2.If you only want to understand the corresponding categories without reading the detailed content, just read the opening paragraph of each major section, such as 1.1 and 1.2, but you don't need to read 1.1.1 or 1.1.2.
It is recommended to consult only the relevant parts of this handbook when needed or to browse only the first-level table of contents; if interested, then browse the full text.It is recommended to consult only the relevant parts of this handbook when needed or to browse only the first-level table of contents; if interested, then browse the full text.
Future updates will include recommended model API service addresses in each section.Future updates will include recommended model API service addresses in each section.
After completing this handbook, you will establish an introductory-level systematic understanding of mainstream AI capabilities — not only knowing "what capabilities are available on the market and which products are commonly paired with them," but also understanding their positions and interrelationships within the overall architecture. You will know how to quickly identify the required capabilities and make informed selections when facing specific business requirements, laying a solid foundation for building AI capability systems.After completing this handbook, you will establish an introductory-level systematic understanding of mainstream AI capabilities — not only knowing "what capabilities are available on the market and which products are commonly paired with them," but also understanding their positions and interrelationships within the overall architecture. You will know how to quickly identify the required capabilities and make informed selections when facing specific business requirements, laying a solid foundation for building AI capability systems.
Before diving into the specific capability map, let's clarify a concept that is frequently mentioned yet somewhat abstract: What exactly counts as a large model? What counts as a small model?Before diving into the specific capability map, let's clarify a concept that is frequently mentioned yet somewhat abstract: What exactly counts as a large model? What counts as a small model?
From an academic perspective, large models typically refer to general-purpose models with parameter counts in the billions, tens of billions, or even trillions, while small models are specialized models tailored for specific tasks or scenarios with smaller parameter counts (tens of millions to hundreds of millions).From an academic perspective, large models typically refer to general-purpose models with parameter counts in the billions, tens of billions, or even trillions, while small models are specialized models tailored for specific tasks or scenarios with smaller parameter counts (tens of millions to hundreds of millions).
From a pricing perspective, if a model's API call is very cheap — for instance, costing a few cents or fractions of a cent per call, or only a few cents per thousand tokens — and there is no particular emphasis on it being a general-purpose large model, then it is typically either a classic small model (e.g., models specifically designed for OCR, ASR, image classification, or content moderation) or a lightweight version of a large model with fewer parameters (compressed or distilled specifically for high concurrency and low cost). If the per-call price is notably higher — say, several dimes or even starting at 1 RMB per call — then it is most likely a large model.From a pricing perspective, if a model's API call is very cheap — for instance, costing a few cents or fractions of a cent per call, or only a few cents per thousand tokens — and there is no particular emphasis on it being a general-purpose large model, then it is typically either a classic small model (e.g., models specifically designed for OCR, ASR, image classification, or content moderation) or a lightweight version of a large model with fewer parameters (compressed or distilled specifically for high concurrency and low cost). If the per-call price is notably higher — say, several dimes or even starting at 1 RMB per call — then it is most likely a large model.
Additionally, if the product copy explicitly emphasizes the use of large language models (LLMs), general-purpose large models, multimodal large models, or mentions completing complex tasks end-to-end from input to output (such as end-to-end conversational bots, end-to-end retrieval Q&A, end-to-end video generation), then it can generally be regarded as a large model.Additionally, if the product copy explicitly emphasizes the use of large language models (LLMs), general-purpose large models, multimodal large models, or mentions completing complex tasks end-to-end from input to output (such as end-to-end conversational bots, end-to-end retrieval Q&A, end-to-end video generation), then it can generally be regarded as a large model.
Conversely, if the promotional focus is on a specific vertical capability — such as bank card recognition, invoice recognition, license plate recognition, ad click-through rate prediction, speech transcription, or content safety moderation — this indicates that the underlying product is more likely one or a group of small models.Conversely, if the promotional focus is on a specific vertical capability — such as bank card recognition, invoice recognition, license plate recognition, ad click-through rate prediction, speech transcription, or content safety moderation — this indicates that the underlying product is more likely one or a group of small models.
Therefore, in the narrative that follows in this article, we can make a pragmatic convention:Therefore, in the narrative that follows in this article, we can make a pragmatic convention:
It's worth supplementing here with a key industry shift: many of the model capabilities mentioned in this handbook were actually handled by "small models" before 2021 — training dedicated models for specific scenarios and specific data to meet precise needs. Today, however, the vast majority of general-purpose scenarios and tasks can already be solved by directly calling large models.It's worth supplementing here with a key industry shift: many of the model capabilities mentioned in this handbook were actually handled by "small models" before 2021 — training dedicated models for specific scenarios and specific data to meet precise needs. Today, however, the vast majority of general-purpose scenarios and tasks can already be solved by directly calling large models.
From the perspective of pursuing the ultimate in precision and cost, the training and application of small models still hold irreplaceable value; but for beginners, we can absolutely start by learning how to find and call large model APIs, then gradually delve into more advanced techniques. You only need to weigh the trade-offs between cost, precision, and latency, then decide where to use general-purpose large models and where to retain or introduce dedicated small models.From the perspective of pursuing the ultimate in precision and cost, the training and application of small models still hold irreplaceable value; but for beginners, we can absolutely start by learning how to find and call large model APIs, then gradually delve into more advanced techniques. You only need to weigh the trade-offs between cost, precision, and latency, then decide where to use general-purpose large models and where to retain or introduce dedicated small models.
> Getting to know common text and multimodal general-purpose large models through some familiar products:> Getting to know common text and multimodal general-purpose large models through some familiar products:
>>
> - OpenAI series: GPT-4, GPT-4.1, GPT-4o, GPT-5.1, etc.> - OpenAI series: GPT-4, GPT-4.1, GPT-4o, GPT-5.1, etc.
> - Google series: Gemini 1.5 Pro, Gemini 1.5 Flash, etc.> - Google series: Gemini 1.5 Pro, Gemini 1.5 Flash, etc.
> - Anthropic series: Claude 3.5 Sonnet, Claude 3.5 Haiku, etc.> - Anthropic series: Claude 3.5 Sonnet, Claude 3.5 Haiku, etc.
> - Domestic models: Tongyi Qianwen (Qwen) series, Wenxin Yiyan (ERNIE Bot) series, GLM/Zhipu Qingyan, Tencent Hunyuan, iFlytek Spark, the large model behind Moonshot AI's Kimi, MiniMax MiniMax-M2.7 series, etc.> - Domestic models: Tongyi Qianwen (Qwen) series, Wenxin Yiyan (ERNIE Bot) series, GLM/Zhipu Qingyan, Tencent Hunyuan, iFlytek Spark, the large model behind Moonshot AI's Kimi, MiniMax MiniMax-M2.7 series, etc.
>>
> Large models and services more oriented toward vision and video include:> Large models and services more oriented toward vision and video include:
>>
> - Image generation: DALL·E, Midjourney, Stable Diffusion, SDXL, Flux, etc.> - Image generation: DALL·E, Midjourney, Stable Diffusion, SDXL, Flux, etc.
> - Multimodal visual understanding: GPT-4o, GPT-4.1 with Vision, Gemini 1.5 (image-text multimodal), Claude 3.5 Sonnet Vision, LLaVA, etc.> - Multimodal visual understanding: GPT-4o, GPT-4.1 with Vision, Gemini 1.5 (image-text multimodal), Claude 3.5 Sonnet Vision, LLaVA, etc.
> - Video generation: Sora, Kling, Runway Gen-2, Pika, Luma, Veo, etc.> - Video generation: Sora, Kling, Runway Gen-2, Pika, Luma, Veo, etc.
>>
> Large models in the voice and audio direction include:> Large models in the voice and audio direction include:
>>
> - Speech recognition ASR: Whisper series (Whisper, Whisper-large-v3, etc.), Deepgram, end-to-end ASR large models from various cloud vendors (such as iFlytek, Baidu, Volcano Engine, Alibaba, etc.)> - Speech recognition ASR: Whisper series (Whisper, Whisper-large-v3, etc.), Deepgram, end-to-end ASR large models from various cloud vendors (such as iFlytek, Baidu, Volcano Engine, Alibaba, etc.)
> - Voice multimodal and voice conversation: GPT-4o (end-to-end voice conversation), OpenAI Realtime, Gemini 1.5's audio understanding capability, etc.> - Voice multimodal and voice conversation: GPT-4o (end-to-end voice conversation), OpenAI Realtime, Gemini 1.5's audio understanding capability, etc.
> - TTS / Audio and music generation: OpenAI TTS, ElevenLabs, Suno, Udio, MusicGen, etc.> - TTS / Audio and music generation: OpenAI TTS, ElevenLabs, Suno, Udio, MusicGen, etc.
>>
> Generation and understanding models in the 3D / spatial direction include:> Generation and understanding models in the 3D / spatial direction include:
>>
> - Text-to-3D and image-to-3D: DreamFusion, Shap-E, GET3D, Zero-1-to-3, TripoSR, etc.> - Text-to-3D and image-to-3D: DreamFusion, Shap-E, GET3D, Zero-1-to-3, TripoSR, etc.
> - NeRF / neural rendering family: Instant-NGP, NeRF series, Gaussian Splatting-related models, etc.> - NeRF / neural rendering family: Instant-NGP, NeRF series, Gaussian Splatting-related models, etc.
Among AI capabilities, text tasks are the most fundamental. Whether we ultimately want to do content moderation, search and recommendation, knowledge Q&A, or writing assistants and code copilots, it all essentially comes down to one question: how to make machines truly understand text.Among AI capabilities, text tasks are the most fundamental. Whether we ultimately want to do content moderation, search and recommendation, knowledge Q&A, or writing assistants and code copilots, it all essentially comes down to one question: how to make machines truly understand text.
Let's start from the most fundamental layer: foundational language modeling and representation. Its role is to first familiarize the machine with language in a statistical sense and, on that basis, find a stable vector/matrix representation for words, sentences, and documents, so as to facilitate downstream tasks such as classification, matching, extraction, and generation. No matter what text-related task you want to do in the future, you will more or less need to answer the same question first: how do I represent this passage with a string of numbers?Let's start from the most fundamental layer: foundational language modeling and representation. Its role is to first familiarize the machine with language in a statistical sense and, on that basis, find a stable vector/matrix representation for words, sentences, and documents, so as to facilitate downstream tasks such as classification, matching, extraction, and generation. No matter what text-related task you want to do in the future, you will more or less need to answer the same question first: how do I represent this passage with a string of numbers?
We can look at this topic from three angles: scenarios, principles, and models:We can look at this topic from three angles: scenarios, principles, and models:
Learning representations of words, sentences, and documents as the foundation for subsequent, more complex tasks.Learning representations of words, sentences, and documents as the foundation for subsequent, more complex tasks.
BERT / RoBERTa / ERNIE, the GPT family, LLaMA / Qwen / Yi and other LLMs; various Embedding models (OpenAI text‑embedding‑3 series, bge, E5, SimCSE, etc.).BERT / RoBERTa / ERNIE, the GPT family, LLaMA / Qwen / Yi and other LLMs; various Embedding models (OpenAI text‑embedding‑3 series, bge, E5, SimCSE, etc.).
The first step at this layer is to let the model become familiar with linguistic patterns through massive amounts of text. The approach can be simply understood as: give the model countless "fill-in-the-blank" exercises — after seeing the context of a passage, have it fill in the most reasonable word (token). With enough practice problems and sufficiently broad corpora, the model gradually learns: what a natural sentence looks like, which words frequently appear together, and what expressions sound awkward. This process is called "language modeling," which is essentially a unified word-guessing training mechanism.The first step at this layer is to let the model become familiar with linguistic patterns through massive amounts of text. The approach can be simply understood as: give the model countless "fill-in-the-blank" exercises — after seeing the context of a passage, have it fill in the most reasonable word (token). With enough practice problems and sufficiently broad corpora, the model gradually learns: what a natural sentence looks like, which words frequently appear together, and what expressions sound awkward. This process is called "language modeling," which is essentially a unified word-guessing training mechanism.
There are two common ways of posing the question, each illustrated with a simple one-sentence example:There are two common ways of posing the question, each illustrated with a simple one-sentence example:
It's raining today, so IInput prefix: It's raining today, so IThis approach primarily trains the model's grasp of continuation, coherence, and common expressions.This approach primarily trains the model's grasp of continuation, coherence, and common expressions.
It's raining today, so I brought an umbrellaOriginal sentence: It's raining today, so I brought an umbrellaToday [MASK] raining, so I brought an umbrellaTraining sentence: Today [MASK] raining, so I brought an umbrella[MASK] with a reasonable word like "is."Model task: fill [MASK] with a reasonable word like "is."Here the model must simultaneously look at the left context "Today" and the right context "so I brought an umbrella" to decide what to fill in, which is more conducive to learning whole-sentence semantics.Here the model must simultaneously look at the left context "Today" and the right context "so I brought an umbrella" to decide what to fill in, which is more conducive to learning whole-sentence semantics.
By repeatedly doing these two types of "word-guessing exercises" on massive corpora, the model gradually accumulates linguistic intuition and statistical common sense. On this foundation, the next step is to explicitly transform this capability into vector representations of words, sentences, and documents, laying the groundwork for subsequent tasks such as retrieval, recommendation, and Q&A.By repeatedly doing these two types of "word-guessing exercises" on massive corpora, the model gradually accumulates linguistic intuition and statistical common sense. On this foundation, the next step is to explicitly transform this capability into vector representations of words, sentences, and documents, laying the groundwork for subsequent tasks such as retrieval, recommendation, and Q&A.
The earliest generation of methods for constructing text vectors was static word vectors: assigning each word a fixed vector that, once trained, does not change with context — intuitive and simple, but unable to distinguish the meanings of polysemous words in different contexts. To solve this problem, context-based dynamic representation methods later emerged: the same word in different sentences generates different vectors, entirely determined by its surrounding context. For example, "apple" in "Apple released a new phone" would lean toward the semantic direction of "tech company," while in "apples are rich in vitamins" it would be closer to the "fruit" concept.The earliest generation of methods for constructing text vectors was static word vectors: assigning each word a fixed vector that, once trained, does not change with context — intuitive and simple, but unable to distinguish the meanings of polysemous words in different contexts. To solve this problem, context-based dynamic representation methods later emerged: the same word in different sentences generates different vectors, entirely determined by its surrounding context. For example, "apple" in "Apple released a new phone" would lean toward the semantic direction of "tech company," while in "apples are rich in vitamins" it would be closer to the "fruit" concept.
This mechanism not only enhances representational capacity at the word level but also paves the way for vectorizing sentences and documents. For sentences, sentence vectors can be generated; for documents, the entire text can be encoded as input (if length permits), or encoded segment by segment and then aggregated into a global vector through attention mechanisms, hierarchical pooling, contrastive learning, or other methods. In recent years, dedicated embedding models (such as bge, E5, and the text-embedding series) have been continuously optimized around the goal of "making semantically similar texts closer in vector space," performing particularly well on tasks such as semantic retrieval and similarity matching.This mechanism not only enhances representational capacity at the word level but also paves the way for vectorizing sentences and documents. For sentences, sentence vectors can be generated; for documents, the entire text can be encoded as input (if length permits), or encoded segment by segment and then aggregated into a global vector through attention mechanisms, hierarchical pooling, contrastive learning, or other methods. In recent years, dedicated embedding models (such as bge, E5, and the text-embedding series) have been continuously optimized around the goal of "making semantically similar texts closer in vector space," performing particularly well on tasks such as semantic retrieval and similarity matching.
This pipeline from contextual modeling to sentence/document vector generation has become the core infrastructure behind systems for search, recommendation, Q&A, and more, bringing us back to the various scenarios mentioned earlier:This pipeline from contextual modeling to sentence/document vector generation has become the core infrastructure behind systems for search, recommendation, Q&A, and more, bringing us back to the various scenarios mentioned earlier:
In engineering terms, the common practice is to encapsulate this into a unified "text vector service": input any piece of text, output a fixed-dimension vector, shared across multiple systems such as search, recommendation, and Q&A. At the product level, the capabilities of this layer are mainly reflected in: semantic recall in search and recommendation (no longer relying solely on keywords, but recalling content that is "worded differently but similar in meaning" through vector similarity), as well as unified embedding/vector retrieval services for enterprise knowledge bases, FAQs, and case libraries.In engineering terms, the common practice is to encapsulate this into a unified "text vector service": input any piece of text, output a fixed-dimension vector, shared across multiple systems such as search, recommendation, and Q&A. At the product level, the capabilities of this layer are mainly reflected in: semantic recall in search and recommendation (no longer relying solely on keywords, but recalling content that is "worded differently but similar in meaning" through vector similarity), as well as unified embedding/vector retrieval services for enterprise knowledge bases, FAQs, and case libraries.
In the previous section, we found the "coordinates" of each piece of text in semantic space through foundational language modeling and representation. But coordinates alone are not enough — the questions that businesses truly care about are often: What category does this text belong to? Is it about the same thing as another piece of text? Do two sentences logically support or contradict each other? You can think of it this way: use the two capabilities of classification and matching to transform the underlying vector representations into labels and relevance signals that can directly drive business decisions. We'll again examine this layer from three angles: scenarios, principles, and models:In the previous section, we found the "coordinates" of each piece of text in semantic space through foundational language modeling and representation. But coordinates alone are not enough — the questions that businesses truly care about are often: What category does this text belong to? Is it about the same thing as another piece of text? Do two sentences logically support or contradict each other? You can think of it this way: use the two capabilities of classification and matching to transform the underlying vector representations into labels and relevance signals that can directly drive business decisions. We'll again examine this layer from three angles: scenarios, principles, and models:
On top of semantic representations, make holistic judgments about an entire piece of text or text pairs:On top of semantic representations, make holistic judgments about an entire piece of text or text pairs:
Based on pre-trained encoders, with simple classification / matching structures attached:Based on pre-trained encoders, with simple classification / matching structures attached:
Leveraging the semantic representations from the previous layer, we can very naturally attach a simple classification head on top and, with a small amount of labeled data, have the model learn to answer one question: "What category does this text belong to?"Leveraging the semantic representations from the previous layer, we can very naturally attach a simple classification head on top and, with a small amount of labeled data, have the model learn to answer one question: "What category does this text belong to?"
The most classic example is sentiment classification. A user's review might be praise, a complaint, or simply a statement of fact. After obtaining the vector representation of the sentence, the model only needs to attach a softmax classification layer to output the probabilities of "positive / negative / neutral." This capability is already very mature in scenarios such as e-commerce, social platforms, and app marketplaces.The most classic example is sentiment classification. A user's review might be praise, a complaint, or simply a statement of fact. After obtaining the vector representation of the sentence, the model only needs to attach a softmax classification layer to output the probabilities of "positive / negative / neutral." This capability is already very mature in scenarios such as e-commerce, social platforms, and app marketplaces.
Another major category is topic / industry classification. In news recommendation, we want to know whether an article is about sports, finance, or entertainment; in an enterprise's internal customer service / ticketing system, the concern is more about whether it's a product inquiry, a functional anomaly, or a complaint/suggestion. These labels can help route content more precisely to the appropriate workflow and can also serve as important features in the recommendation and ranking stage.Another major category is topic / industry classification. In news recommendation, we want to know whether an article is about sports, finance, or entertainment; in an enterprise's internal customer service / ticketing system, the concern is more about whether it's a product inquiry, a functional anomaly, or a complaint/suggestion. These labels can help route content more precisely to the appropriate workflow and can also serve as important features in the recommendation and ranking stage.
Going further, risk / compliance classification is directly tied to platform safety. We set up dedicated classification models for categories such as ad trafficking, abusive attacks, politically sensitive content, and vulgar/pornographic material, working in conjunction with human review to intercept or downgrade high-risk content. It can be said that the first gate of the vast majority of content safety strategies is built from these types of classifiers.Going further, risk / compliance classification is directly tied to platform safety. We set up dedicated classification models for categories such as ad trafficking, abusive attacks, politically sensitive content, and vulgar/pornographic material, working in conjunction with human review to intercept or downgrade high-risk content. It can be said that the first gate of the vast majority of content safety strategies is built from these types of classifiers.
As we can see, by this layer we are already able to transform "abstract semantic representations" into several business-usable labels. Next, we will discuss: when relationships arise between pieces of text, how do we perform matching and inference.As we can see, by this layer we are already able to transform "abstract semantic representations" into several business-usable labels. Next, we will discuss: when relationships arise between pieces of text, how do we perform matching and inference.
Unlike classification, which "characterizes individual pieces of text," text matching focuses on "the relevance between two pieces of text." In many products, this is often the key link in achieving "intelligence": whether the system can find the most suitable response in the knowledge base when a user says something depends entirely on matching quality.Unlike classification, which "characterizes individual pieces of text," text matching focuses on "the relevance between two pieces of text." In many products, this is often the key link in achieving "intelligence": whether the system can find the most suitable response in the knowledge base when a user says something depends entirely on matching quality.
The most fundamental aspect is semantic similarity calculation. We first use the embedding model from the previous layer to encode two sentences into vectors, then determine their distance in semantic space through cosine similarity, dot product, or other methods. Models like SimCSE and Sentence‑BERT are specifically designed, through contrastive learning, to pull "similar sentence pairs" closer and push "dissimilar sentence pairs" farther apart.The most fundamental aspect is semantic similarity calculation. We first use the embedding model from the previous layer to encode two sentences into vectors, then determine their distance in semantic space through cosine similarity, dot product, or other methods. Models like SimCSE and Sentence‑BERT are specifically designed, through contrastive learning, to pull "similar sentence pairs" closer and push "dissimilar sentence pairs" farther apart.
Building on this, paraphrase detection and plagiarism detection are simply matching tasks for specific application scenarios. The former is used for content deduplication, preventing platforms from being flooded with repetitive expressions; the latter is used in scenarios such as education and knowledge communities to identify highly similar answers or articles. Technically, both essentially involve binary classification or ranking based on text similarity.Building on this, paraphrase detection and plagiarism detection are simply matching tasks for specific application scenarios. The former is used for content deduplication, preventing platforms from being flooded with repetitive expressions; the latter is used in scenarios such as education and knowledge communities to identify highly similar answers or articles. Technically, both essentially involve binary classification or ranking based on text similarity.
A very important downstream application is Q&A matching. When a user poses a natural language question, we don't directly use keywords to match FAQs; instead, we first perform recall through semantic vectors, then use a finer-grained matching model (such as a Cross‑Encoder) to re-rank several candidates and select the most likely corresponding one. This pipeline forms the foundation of FAQ bots and document Q&A systems.A very important downstream application is Q&A matching. When a user poses a natural language question, we don't directly use keywords to match FAQs; instead, we first perform recall through semantic vectors, then use a finer-grained matching model (such as a Cross‑Encoder) to re-rank several candidates and select the most likely corresponding one. This pipeline forms the foundation of FAQ bots and document Q&A systems.
At this layer, we already have the ability to classify and judge relationships for "entire pieces of text." But in many scenarios, businesses are not satisfied with this alone — they further want to know: what specific entities are mentioned in this text, and what events occurred. This naturally leads to the topic of the next section — sequence labeling and information extraction.At this layer, we already have the ability to classify and judge relationships for "entire pieces of text." But in many scenarios, businesses are not satisfied with this alone — they further want to know: what specific entities are mentioned in this text, and what events occurred. This naturally leads to the topic of the next section — sequence labeling and information extraction.
After completing the classification and matching of entire texts, we often encounter a more granular requirement: not only knowing "what this article is about and how risky it is," but also knowing "who it specifically mentions, where, when, and what the amount is." This section is a key step beyond holistic judgment toward "fine-grained structuring." You can think of it as: on the premise of already knowing "which category of text to look at and roughly what it's about," we mine entities, relations, events, and various fields from within the text, enabling unstructured text to be directly consumed by business systems. We'll again look at this layer from four aspects: scenarios, principles, models, and products:After completing the classification and matching of entire texts, we often encounter a more granular requirement: not only knowing "what this article is about and how risky it is," but also knowing "who it specifically mentions, where, when, and what the amount is." This section is a key step beyond holistic judgment toward "fine-grained structuring." You can think of it as: on the premise of already knowing "which category of text to look at and roughly what it's about," we mine entities, relations, events, and various fields from within the text, enabling unstructured text to be directly consumed by business systems. We'll again look at this layer from four aspects: scenarios, principles, models, and products:
Perform fine-grained labeling and structuring of text at the token / phrase level:Perform fine-grained labeling and structuring of text at the token / phrase level:
Based on pre-trained representations, perform information extraction through structures such as sequence labeling or span extraction:Based on pre-trained representations, perform information extraction through structures such as sequence labeling or span extraction:
In the text classification stage, we only cared about what category an entire piece of text belongs to; in the sequence labeling stage, we need to label every token and every phrase in the text. The most typical task is Named Entity Recognition (NER): identifying specific types of entities such as person names, organization names, location names, product names, and disease names.In the text classification stage, we only cared about what category an entire piece of text belongs to; in the sequence labeling stage, we need to label every token and every phrase in the text. The most typical task is Named Entity Recognition (NER): identifying specific types of entities such as person names, organization names, location names, product names, and disease names.
In terms of modeling approaches, the traditional method uses sequence labeling structures like BiLSTM + CRF, while later approaches more commonly adopt BERT + CRF or BERT + Softmax, leveraging the contextual representation capability of pre-trained encoders to determine each token's label (e.g., B‑ORG, I‑ORG, O, etc.). In practice, NER models are often the first "preprocessing" step for subsequent knowledge graphs and relation extraction.In terms of modeling approaches, the traditional method uses sequence labeling structures like BiLSTM + CRF, while later approaches more commonly adopt BERT + CRF or BERT + Softmax, leveraging the contextual representation capability of pre-trained encoders to determine each token's label (e.g., B‑ORG, I‑ORG, O, etc.). In practice, NER models are often the first "preprocessing" step for subsequent knowledge graphs and relation extraction.
Beyond NER, part-of-speech tagging and phrase segmentation are also typical sequence labeling tasks. They serve more for underlying linguistic analysis, providing basic structures for subsequent, more complex syntactic / semantic tasks.Beyond NER, part-of-speech tagging and phrase segmentation are also typical sequence labeling tasks. They serve more for underlying linguistic analysis, providing basic structures for subsequent, more complex syntactic / semantic tasks.
Once we have identified entities in the text through sequence labeling, a natural next question is: what exactly are the relationships between these entities, and what kind of events do they collectively constitute?Once we have identified entities in the text through sequence labeling, a natural next question is: what exactly are the relationships between these entities, and what kind of events do they collectively constitute?
Relation extraction focuses on "entity pairs + relation types." For example, in the sentence "Zhang San joined a certain tech company as CTO in 2024," we need to not only identify the two entities "Zhang San" and "a certain tech company" but also extract the "employed at" relationship between them.Relation extraction focuses on "entity pairs + relation types." For example, in the sentence "Zhang San joined a certain tech company as CTO in 2024," we need to not only identify the two entities "Zhang San" and "a certain tech company" but also extract the "employed at" relationship between them.
Building on relations, event extraction attempts to reconstruct "who did what, when, and where." Taking a news article as an example, a standard event template might include multiple slots: event type (acquisition, partnership, accident), time, location, participants, amount, consequences, etc. Event extraction models need to automatically fill these slots from lengthy text, thereby constructing an "event table" that can be retrieved, statistically analyzed, and reasoned about.Building on relations, event extraction attempts to reconstruct "who did what, when, and where." Taking a news article as an example, a standard event template might include multiple slots: event type (acquisition, partnership, accident), time, location, participants, amount, consequences, etc. Event extraction models need to automatically fill these slots from lengthy text, thereby constructing an "event table" that can be retrieved, statistically analyzed, and reasoned about.
In terms of modeling methods, beyond traditional sequence-labeling-style extraction, we also use Span‑based IE (directly predicting the start and end positions of entity / relation spans) as well as the more recently emerged Prompt‑based IE and LLM-based Few‑shot extraction. The advantage of the latter is that new schemas can be quickly adapted through natural language prompts, significantly reducing the cost of extensive re-labeling and retraining.In terms of modeling methods, beyond traditional sequence-labeling-style extraction, we also use Span‑based IE (directly predicting the start and end positions of entity / relation spans) as well as the more recently emerged Prompt‑based IE and LLM-based Few‑shot extraction. The advantage of the latter is that new schemas can be quickly adapted through natural language prompts, significantly reducing the cost of extensive re-labeling and retraining.
From an engineering perspective, a mature extraction system typically forms a pipeline:From an engineering perspective, a mature extraction system typically forms a pipeline:
In the preceding sections, we sequentially built the understanding pipeline of "representation → classification and matching → sequence labeling and extraction": the model can not only map text into semantic space but also make judgments about entire pieces of text and extract structured information from them. What this section aims to do is "reverse" that understanding pipeline: on the basis of sufficient understanding, have the model proactively produce, rewrite, compress, and polish text. You can think of it as: performing "reverse encoding" in semantic space, turning internal representations back into high-quality natural language output — the layer closest to user perception in the entire text modality capability chain. We'll again break it down from four dimensions: scenarios, principles, models, and products:In the preceding sections, we sequentially built the understanding pipeline of "representation → classification and matching → sequence labeling and extraction": the model can not only map text into semantic space but also make judgments about entire pieces of text and extract structured information from them. What this section aims to do is "reverse" that understanding pipeline: on the basis of sufficient understanding, have the model proactively produce, rewrite, compress, and polish text. You can think of it as: performing "reverse encoding" in semantic space, turning internal representations back into high-quality natural language output — the layer closest to user perception in the entire text modality capability chain. We'll again break it down from four dimensions: scenarios, principles, models, and products:
On the foundation of language modeling, perform "creation from scratch" and "modification based on existing content" on text:On the foundation of language modeling, perform "creation from scratch" and "modification based on existing content" on text:
Primarily large-scale pre-trained + instruction fine-tuned generative models:Primarily large-scale pre-trained + instruction fine-tuned generative models:
Since this section is essentially equivalent to prompt engineering, it will not be elaborated further here; you can refer to the prompt engineering tutorial section on your own.Since this section is essentially equivalent to prompt engineering, it will not be elaborated further here; you can refer to the prompt engineering tutorial section on your own.
In AI capabilities, the image modality is responsible for "understanding the world through vision." Whether the goal is security surveillance, autonomous driving, short video effects, e-commerce intelligent retouching, multimodal Q&A, or AI painting, it all essentially follows one path: starting from raw pixels and progressively obtaining structured understanding and controllable generation of images.In AI capabilities, the image modality is responsible for "understanding the world through vision." Whether the goal is security surveillance, autonomous driving, short video effects, e-commerce intelligent retouching, multimodal Q&A, or AI painting, it all essentially follows one path: starting from raw pixels and progressively obtaining structured understanding and controllable generation of images.
In the previous section, we introduced the role of the vision modality in multimodal systems and how it interfaces with language and speech. But before diving into "high-level semantic tasks" like object detection, image understanding, and visual question answering, there is a foundational capability layer that is often overlooked yet critically important — low-level vision. You can think of it this way: before "understanding what is in the image," the system needs to first address two questions: "how good is the quality of this image itself" and "what stable local structures can be reused by upper layers." Through a layer of universal restoration, enhancement, and structure extraction, raw pixels are transformed into cleaner, more stable image representations.In the previous section, we introduced the role of the vision modality in multimodal systems and how it interfaces with language and speech. But before diving into "high-level semantic tasks" like object detection, image understanding, and visual question answering, there is a foundational capability layer that is often overlooked yet critically important — low-level vision. You can think of it this way: before "understanding what is in the image," the system needs to first address two questions: "how good is the quality of this image itself" and "what stable local structures can be reused by upper layers." Through a layer of universal restoration, enhancement, and structure extraction, raw pixels are transformed into cleaner, more stable image representations.
From an engineering perspective, low-level vision directly affects both the "image quality experience" perceived by users and the health of input distributions for downstream tasks like detection, recognition, and segmentation. If this layer is not done well, all subsequent models have to struggle in environments with "heavy noise, severe distortion, and extreme lighting." Conversely, if images are repaired as much as possible and structural information is well extracted at this layer, high-level tasks can perform on a much friendlier foundation. Below, we organize this layer from three angles: scenarios, principles, and models.From an engineering perspective, low-level vision directly affects both the "image quality experience" perceived by users and the health of input distributions for downstream tasks like detection, recognition, and segmentation. If this layer is not done well, all subsequent models have to struggle in environments with "heavy noise, severe distortion, and extreme lighting." Conversely, if images are repaired as much as possible and structural information is well extracted at this layer, high-level tasks can perform on a much friendlier foundation. Below, we organize this layer from three angles: scenarios, principles, and models.
Centered on two core objectives — "image quality" and "local structure" — physical and statistical modeling is applied to pixel-level information:Centered on two core objectives — "image quality" and "local structure" — physical and statistical modeling is applied to pixel-level information:
A combination of classical image processing methods and deep learning models, balancing efficiency and effectiveness:A combination of classical image processing methods and deep learning models, balancing efficiency and effectiveness:
In low-level vision, image restoration and enhancement first confront various types of degradation: noise, blur, compression artifacts, low light, insufficient dynamic range, etc. Raw images in many real-world scenarios are not "clean": night scenes and indoor low light fill the frame with grain and color noise; snapshots and surveillance footage often appear blurry due to motion or poor focus; video compression introduces blocky artifacts. The goal of restoration and enhancement is to restore clear details and natural visual perception as much as possible without altering the semantic content of the image — turning "blurry, dark, dirty" inputs into "clear, bright, comfortable" ones.In low-level vision, image restoration and enhancement first confront various types of degradation: noise, blur, compression artifacts, low light, insufficient dynamic range, etc. Raw images in many real-world scenarios are not "clean": night scenes and indoor low light fill the frame with grain and color noise; snapshots and surveillance footage often appear blurry due to motion or poor focus; video compression introduces blocky artifacts. The goal of restoration and enhancement is to restore clear details and natural visual perception as much as possible without altering the semantic content of the image — turning "blurry, dark, dirty" inputs into "clear, bright, comfortable" ones.
Typical tasks include denoising, deblurring, low-light enhancement, and super-resolution. Denoising and deblurring require a trade-off between local texture and global structure: suppressing high-frequency noise and deconvolving blur kernels without also smoothing away real details. Low-light enhancement must boost brightness and contrast while avoiding amplifying dark-area noise, correcting color casts, and controlling overexposed regions. Super-resolution focuses on adding reasonable high-frequency information during upscaling, so the enlarged image neither looks "blurry and plasticky" nor excessively "fabricates" details. Modern approaches mostly use deep networks (CNN or Vision Transformer), learning the mapping from observed image y to ideal image x on large-scale "degraded–clean" paired data, with composite objectives including pixel error, perceptual loss, and adversarial loss to strike a balance between "good metrics" and "pleasing to the human eye."Typical tasks include denoising, deblurring, low-light enhancement, and super-resolution. Denoising and deblurring require a trade-off between local texture and global structure: suppressing high-frequency noise and deconvolving blur kernels without also smoothing away real details. Low-light enhancement must boost brightness and contrast while avoiding amplifying dark-area noise, correcting color casts, and controlling overexposed regions. Super-resolution focuses on adding reasonable high-frequency information during upscaling, so the enlarged image neither looks "blurry and plasticky" nor excessively "fabricates" details. Modern approaches mostly use deep networks (CNN or Vision Transformer), learning the mapping from observed image y to ideal image x on large-scale "degraded–clean" paired data, with composite objectives including pixel error, perceptual loss, and adversarial loss to strike a balance between "good metrics" and "pleasing to the human eye."
These capabilities often manifest implicitly in products: the night mode and HDR capture on phone cameras, one-click quality enhancement on short video platforms, old photo restoration tools, and cloud-based enhancement services for surveillance systems all fundamentally rely on this layer's restoration and enhancement modules. For businesses, they directly affect users' subjective perception of "image quality" and indirectly determine the input quality for downstream detection, recognition, and segmentation algorithms. One could say that the more complex the upper-level vision tasks, the more they depend on a high-quality, distributionally stable "image foundation" at the bottom layer.These capabilities often manifest implicitly in products: the night mode and HDR capture on phone cameras, one-click quality enhancement on short video platforms, old photo restoration tools, and cloud-based enhancement services for surveillance systems all fundamentally rely on this layer's restoration and enhancement modules. For businesses, they directly affect users' subjective perception of "image quality" and indirectly determine the input quality for downstream detection, recognition, and segmentation algorithms. One could say that the more complex the upper-level vision tasks, the more they depend on a high-quality, distributionally stable "image foundation" at the bottom layer.
Once image quality has been restored to a usable level, the second key task of low-level vision is to extract features from pixels that are temporarily unrelated to specific semantics but are very important for geometric structure and visual perception, and to unify geometry and illumination. This step won't directly tell you "this is a car" or "this is someone's face," but it will answer questions like "where are the clear contours and corners," "which regions have significant texture structure," and "is the image distorted or tilted," providing reliable structural input for upper-level models.Once image quality has been restored to a usable level, the second key task of low-level vision is to extract features from pixels that are temporarily unrelated to specific semantics but are very important for geometric structure and visual perception, and to unify geometry and illumination. This step won't directly tell you "this is a car" or "this is someone's face," but it will answer questions like "where are the clear contours and corners," "which regions have significant texture structure," and "is the image distorted or tilted," providing reliable structural input for upper-level models.
In terms of feature extraction, edges and corners are the most fundamental elements. Using operators like Canny and Sobel, the system can mark the "edges" where grayscale or color changes most sharply across the entire image — these often correspond to object contours, component boundaries, and texture directions. Corner detection (e.g., Harris, FAST) finds "corners" where local gradients change significantly in multiple directions, typically appearing at object corners and line intersections. Further, local descriptors like SIFT, SURF, and ORB encode the texture pattern of a small region around these keypoints, enabling the same physical point to be matched across different viewpoints, scales, and certain illumination changes. This provides foundational support for image registration, panorama stitching, SLAM, AR tracking, and 3D reconstruction.In terms of feature extraction, edges and corners are the most fundamental elements. Using operators like Canny and Sobel, the system can mark the "edges" where grayscale or color changes most sharply across the entire image — these often correspond to object contours, component boundaries, and texture directions. Corner detection (e.g., Harris, FAST) finds "corners" where local gradients change significantly in multiple directions, typically appearing at object corners and line intersections. Further, local descriptors like SIFT, SURF, and ORB encode the texture pattern of a small region around these keypoints, enabling the same physical point to be matched across different viewpoints, scales, and certain illumination changes. This provides foundational support for image registration, panorama stitching, SLAM, AR tracking, and 3D reconstruction.
In parallel with feature extraction are various geometric and illumination preprocessing operations. Barrel/pincushion distortion from wide-angle lenses, and skew and perspective stretching when photographing documents, are identified through low-level geometric cues like line detection and vanishing point estimation, and are "pulled back to normal" through undistortion, straightening, and perspective correction. Global or adaptive histogram equalization, contrast stretching, and illumination normalization enhance local contrast and reduce the impact of uneven lighting and shadows while preserving details. Color space transformations (RGB → HSV/Lab) and color histogram statistics provide directly usable input for simple color-based segmentation, saliency region detection, and color cast correction.In parallel with feature extraction are various geometric and illumination preprocessing operations. Barrel/pincushion distortion from wide-angle lenses, and skew and perspective stretching when photographing documents, are identified through low-level geometric cues like line detection and vanishing point estimation, and are "pulled back to normal" through undistortion, straightening, and perspective correction. Global or adaptive histogram equalization, contrast stretching, and illumination normalization enhance local contrast and reduce the impact of uneven lighting and shadows while preserving details. Color space transformations (RGB → HSV/Lab) and color histogram statistics provide directly usable input for simple color-based segmentation, saliency region detection, and color cast correction.
After end-to-end deep learning became mainstream, some of these structural features and preprocessing steps have been "internalized" into the convolutional kernels and normalization strategies of the first few network layers, no longer appearing as explicit operators in system architecture diagrams. But functionally, they still play the same role: first using a relatively universal, class-agnostic low-level processing layer to organize raw pixels into representations that are more stable in terms of geometric shape, lighting conditions, and local structure, then handing them off to upper-level classification, detection, segmentation, and multimodal modules to complete the task of "understanding what this is." Without this layer of "scaffolding," upper-level models would have to struggle on raw images with heavy noise, severe distortion, and blurred structures, and the overall system's robustness and generalization ability would significantly degrade.After end-to-end deep learning became mainstream, some of these structural features and preprocessing steps have been "internalized" into the convolutional kernels and normalization strategies of the first few network layers, no longer appearing as explicit operators in system architecture diagrams. But functionally, they still play the same role: first using a relatively universal, class-agnostic low-level processing layer to organize raw pixels into representations that are more stable in terms of geometric shape, lighting conditions, and local structure, then handing them off to upper-level classification, detection, segmentation, and multimodal modules to complete the task of "understanding what this is." Without this layer of "scaffolding," upper-level models would have to struggle on raw images with heavy noise, severe distortion, and blurred structures, and the overall system's robustness and generalization ability would significantly degrade.
In most image tasks, the questions businesses truly care about are: What category does this entire image belong to? Who is the person in this image? Is this pedestrian the same person across different cameras? You can think of this layer as: on a unified, clean input space, assigning "category labels" or "identity labels" to entire images or entire persons/objects, converting visual signals into the most directly usable recognition results.In most image tasks, the questions businesses truly care about are: What category does this entire image belong to? Who is the person in this image? Is this pedestrian the same person across different cameras? You can think of this layer as: on a unified, clean input space, assigning "category labels" or "identity labels" to entire images or entire persons/objects, converting visual signals into the most directly usable recognition results.
From a product perspective, image classification and recognition were among the first vision capabilities to be deployed at scale and serve as the "entry module" for many upper-level applications. E-commerce and content platforms use it to automatically tag images and identify main product categories; security and access control systems use it to confirm "is this the same person"; person re-identification systems sift through feeds from multiple cameras to find the cross-scene trajectory of the same target. Below, we organize this layer from three angles: scenarios, principles, and models.From a product perspective, image classification and recognition were among the first vision capabilities to be deployed at scale and serve as the "entry module" for many upper-level applications. E-commerce and content platforms use it to automatically tag images and identify main product categories; security and access control systems use it to confirm "is this the same person"; person re-identification systems sift through feeds from multiple cameras to find the cross-scene trajectory of the same target. Below, we organize this layer from three angles: scenarios, principles, and models.
Discriminative modeling of entire images or entire persons/objects in a unified visual feature space:Discriminative modeling of entire images or entire persons/objects in a unified visual feature space:
Deep convolutional networks and Vision Transformers serve as the backbone, combined with classification heads or metric learning heads to implement different types of recognition tasks:Deep convolutional networks and Vision Transformers serve as the backbone, combined with classification heads or metric learning heads to implement different types of recognition tasks:
Corresponding to specific product forms, the capabilities of this layer are often delivered as "image content recognition / classification APIs," "face recognition SDK / SaaS," "person re-identification platforms," etc. They often directly drive business decisions (e.g., access control clearance, content tag writing) and also serve as upstream, providing structured labels and stable identity representations for subsequent retrieval, recommendation, behavior analysis, and multimodal understanding. Below, we expand on two directions: image classification and identity/attribute recognition.Corresponding to specific product forms, the capabilities of this layer are often delivered as "image content recognition / classification APIs," "face recognition SDK / SaaS," "person re-identification platforms," etc. They often directly drive business decisions (e.g., access control clearance, content tag writing) and also serve as upstream, providing structured labels and stable identity representations for subsequent retrieval, recommendation, behavior analysis, and multimodal understanding. Below, we expand on two directions: image classification and identity/attribute recognition.
In the most basic image classification task, the system faces an entire image and the goal is to assign it one or several semantic category labels. The most common form is single-label classification — for example, in datasets like ImageNet, each image is labeled with one primary category such as "dog," "cat," "car," "airplane." In business scenarios, this capability is widely used to tag user-uploaded images with theme labels like "landscape / food / pet / portrait / document" to support retrieval, recommendation, and content moderation. Similar to text classification, the model attaches a fully connected + Softmax layer on top of global visual features extracted by a pre-trained backbone, outputting a probability distribution over all candidate categories.In the most basic image classification task, the system faces an entire image and the goal is to assign it one or several semantic category labels. The most common form is single-label classification — for example, in datasets like ImageNet, each image is labeled with one primary category such as "dog," "cat," "car," "airplane." In business scenarios, this capability is widely used to tag user-uploaded images with theme labels like "landscape / food / pet / portrait / document" to support retrieval, recommendation, and content moderation. Similar to text classification, the model attaches a fully connected + Softmax layer on top of global visual features extracted by a pre-trained backbone, outputting a probability distribution over all candidate categories.
In many real-world applications, an image often belongs to multiple categories simultaneously. For example, a "seaside sunset selfie" image could be "landscape," "portrait," and also tagged as "travel" or "seaside." This calls for multi-label classification: the model still starts from whole-image features, but the output layer is no longer a mutually exclusive Softmax — instead, it predicts the presence/absence probability (Sigmoid) for each label independently, trained with a multi-label loss function. To handle the large number of "long-tail categories" (rare labels with very few samples) in real-world data, multi-label classification models often incorporate class re-weighting, hard example mining, or label structure modeling to improve recall for niche categories.In many real-world applications, an image often belongs to multiple categories simultaneously. For example, a "seaside sunset selfie" image could be "landscape," "portrait," and also tagged as "travel" or "seaside." This calls for multi-label classification: the model still starts from whole-image features, but the output layer is no longer a mutually exclusive Softmax — instead, it predicts the presence/absence probability (Sigmoid) for each label independently, trained with a multi-label loss function. To handle the large number of "long-tail categories" (rare labels with very few samples) in real-world data, multi-label classification models often incorporate class re-weighting, hard example mining, or label structure modeling to improve recall for niche categories.
At the human-machine interface level, image classification is typically provided as an "image content recognition API." Upstream businesses only need to upload an image to receive a set of category labels and their confidence scores for subsequent decision-making: for example, an ad delivery system can restrict certain sensitive categories based on image content; an e-commerce platform can use image classification to assist with product category correction; a content platform can use it to enrich recommendation features and moderation signals. Although technically relatively mature, this capability remains the cornerstone for more complex downstream capabilities like object detection, instance segmentation, and visual question answering.At the human-machine interface level, image classification is typically provided as an "image content recognition API." Upstream businesses only need to upload an image to receive a set of category labels and their confidence scores for subsequent decision-making: for example, an ad delivery system can restrict certain sensitive categories based on image content; an e-commerce platform can use image classification to assist with product category correction; a content platform can use it to enrich recommendation features and moderation signals. Although technically relatively mature, this capability remains the cornerstone for more complex downstream capabilities like object detection, instance segmentation, and visual question answering.
Unlike "what type of image is this," image recognition is more concerned with "who is the person/object in this image" — that is, identity-level, instance-level differentiation. Typical representatives are face recognition and person re-identification: the former determines "which identity in the gallery is closest to the current face" in scenarios like access control, attendance, and payment; the latter searches across multiple camera feeds and time periods to find whether the same pedestrian appears, assisting with case review and trajectory analysis. The core of these tasks is no longer simple multi-class classification, but rather learning an embedding in feature space that is "compact within classes and separated between classes," so that images of the same identity captured under different poses, lighting, and cameras can still be clustered together.Unlike "what type of image is this," image recognition is more concerned with "who is the person/object in this image" — that is, identity-level, instance-level differentiation. Typical representatives are face recognition and person re-identification: the former determines "which identity in the gallery is closest to the current face" in scenarios like access control, attendance, and payment; the latter searches across multiple camera feeds and time periods to find whether the same pedestrian appears, assisting with case review and trajectory analysis. The core of these tasks is no longer simple multi-class classification, but rather learning an embedding in feature space that is "compact within classes and separated between classes," so that images of the same identity captured under different poses, lighting, and cameras can still be clustered together.
In model design, face recognition and person re-identification typically adopt a similar paradigm: first use backbones like ResNet, ConvNeXt, ViT, Swin to extract face/person-centered features, then apply loss functions specifically designed for metric learning, such as ArcFace, CosFace, etc. Unlike ordinary classification losses, these losses directly constrain inter-class boundaries in angular space or feature space, explicitly widening the margin between features of different identities, so that the trained features can be used for large-scale vector retrieval without being limited to the fixed classes seen during training. During online serving, the system pre-computes and indexes features for each identity in the gallery, then performs approximate nearest neighbor search on the query face/person features to find the most similar candidates, making final decisions based on business thresholds and multimodal information.In model design, face recognition and person re-identification typically adopt a similar paradigm: first use backbones like ResNet, ConvNeXt, ViT, Swin to extract face/person-centered features, then apply loss functions specifically designed for metric learning, such as ArcFace, CosFace, etc. Unlike ordinary classification losses, these losses directly constrain inter-class boundaries in angular space or feature space, explicitly widening the margin between features of different identities, so that the trained features can be used for large-scale vector retrieval without being limited to the fixed classes seen during training. During online serving, the system pre-computes and indexes features for each identity in the gallery, then performs approximate nearest neighbor search on the query face/person features to find the most similar candidates, making final decisions based on business thresholds and multimodal information.
Complementing "direct identity recognition" is attribute recognition, which does not point to a specific person. In many security and retail scenarios, the system only needs to know attributes like "male or female," "approximate age group," "whether wearing a hat/mask," "clothing color and style," "whether carrying a backpack/luggage," etc., for quick target filtering — without, and often without being appropriate to, directly outputting personal identity. These tasks typically attach multiple parallel attribute heads on top of shared pedestrian/human features (a "head" here means the position where probabilities are output; multiple probability output positions can be used to determine categories), with each head responsible for predicting one or a group of attribute labels, forming a multi-task learning framework. On one hand, multi-task training can make features richer and better generalized; on the other hand, attributes themselves can serve as auxiliary conditions for Re-ID or retrieval, improving system usability in complex scenarios.Complementing "direct identity recognition" is attribute recognition, which does not point to a specific person. In many security and retail scenarios, the system only needs to know attributes like "male or female," "approximate age group," "whether wearing a hat/mask," "clothing color and style," "whether carrying a backpack/luggage," etc., for quick target filtering — without, and often without being appropriate to, directly outputting personal identity. These tasks typically attach multiple parallel attribute heads on top of shared pedestrian/human features (a "head" here means the position where probabilities are output; multiple probability output positions can be used to determine categories), with each head responsible for predicting one or a group of attribute labels, forming a multi-task learning framework. On one hand, multi-task training can make features richer and better generalized; on the other hand, attributes themselves can serve as auxiliary conditions for Re-ID or retrieval, improving system usability in complex scenarios.
In product form, these capabilities are typically packaged as "face recognition SDK / cloud service," "person re-identification platform," "human attribute recognition API," etc., and integrated into access control gates, attendance machines, security platforms, and video structuring systems. Compared to general image classification, they have higher requirements for data security and privacy protection, and are more sensitive to the trade-off between false recognition rate and recall rate. Therefore, beyond algorithms, they are supplemented with mechanisms such as quality detection (e.g., whether it is a real person, whether occluded/recaptured), liveness detection, and multimodal cross-verification, forming a more complete and responsible identity recognition solution.In product form, these capabilities are typically packaged as "face recognition SDK / cloud service," "person re-identification platform," "human attribute recognition API," etc., and integrated into access control gates, attendance machines, security platforms, and video structuring systems. Compared to general image classification, they have higher requirements for data security and privacy protection, and are more sensitive to the trade-off between false recognition rate and recall rate. Therefore, beyond algorithms, they are supplemented with mechanisms such as quality detection (e.g., whether it is a real person, whether occluded/recaptured), liveness detection, and multimodal cross-verification, forming a more complete and responsible identity recognition solution.
In the previous image classification and recognition sections, we only assigned an overall label to "the entire image" or "the entire person," ignoring where and at what size the object appears in the image. However, a more common question in real business is: What objects are in this image? Where is each one located? For example, in a street scene image, we want to simultaneously mark all pedestrians, vehicles, and traffic signs; on an industrial production line, we need to mark all defect areas and part positions in the same frame. Object detection is designed for these needs: in a single image or video frame, it simultaneously predicts the position (bounding box) and category of every object, serving as the foundational capability for many downstream vision tasks (tracking, segmentation, behavior analysis, multi-object counting, etc.).In the previous image classification and recognition sections, we only assigned an overall label to "the entire image" or "the entire person," ignoring where and at what size the object appears in the image. However, a more common question in real business is: What objects are in this image? Where is each one located? For example, in a street scene image, we want to simultaneously mark all pedestrians, vehicles, and traffic signs; on an industrial production line, we need to mark all defect areas and part positions in the same frame. Object detection is designed for these needs: in a single image or video frame, it simultaneously predicts the position (bounding box) and category of every object, serving as the foundational capability for many downstream vision tasks (tracking, segmentation, behavior analysis, multi-object counting, etc.).
From an engineering usage perspective, object detection is the "first step of structuring" for many vision systems — decomposing a raw image into several labeled rectangular boxes, each of which can be further sent to other modules for recognition, tracking, attribute analysis, and even semantic generation. Pedestrian/vehicle detection in surveillance cameras, product detection on unmanned retail shelves, defect/foreign object detection in industrial quality inspection, and the "object detection" APIs provided by cloud vendors all fundamentally rely on this layer of capability. Below, we organize object detection from three angles — scenarios, principles, and models — and expand on key directions in subsequent subsections.From an engineering usage perspective, object detection is the "first step of structuring" for many vision systems — decomposing a raw image into several labeled rectangular boxes, each of which can be further sent to other modules for recognition, tracking, attribute analysis, and even semantic generation. Pedestrian/vehicle detection in surveillance cameras, product detection on unmanned retail shelves, defect/foreign object detection in industrial quality inspection, and the "object detection" APIs provided by cloud vendors all fundamentally rely on this layer of capability. Below, we organize object detection from three angles — scenarios, principles, and models — and expand on key directions in subsequent subsections.
The core of object detection is building a dense prediction mechanism on the image:The core of object detection is building a dense prediction mechanism on the image:
Detection models are generally composed of three parts: backbone network + feature pyramid / head structure + loss and post-processing:Detection models are generally composed of three parts: backbone network + feature pyramid / head structure + loss and post-processing:
Overall, object detection sits at the "hub position" of the vision capability spectrum — on one hand, it receives clean image inputs provided by low-level vision; on the other hand, it decomposes images into "object-level" elements that can be used for recognition, tracking, segmentation, and multimodal understanding. Below, we expand on three directions: single/two-stage detection architectures, anchor-based / anchor-free / Transformer detection, and small object and video detection.Overall, object detection sits at the "hub position" of the vision capability spectrum — on one hand, it receives clean image inputs provided by low-level vision; on the other hand, it decomposes images into "object-level" elements that can be used for recognition, tracking, segmentation, and multimodal understanding. Below, we expand on three directions: single/two-stage detection architectures, anchor-based / anchor-free / Transformer detection, and small object and video detection.
From an architectural perspective, the most classic division in object detection is two-stage vs. one-stage. The main difference between the two is: whether to first "roughly select a batch of candidate boxes and then refine them," or to "predict all boxes and categories in one go on the feature map."From an architectural perspective, the most classic division in object detection is two-stage vs. one-stage. The main difference between the two is: whether to first "roughly select a batch of candidate boxes and then refine them," or to "predict all boxes and categories in one go on the feature map."
Two-stage detection is represented by Faster R-CNN. It first generates a batch of candidate boxes with "high probability of containing an object" through an RPN (Region Proposal Network) on the backbone feature maps (first stage), then performs RoI alignment and feature extraction for each candidate region, followed by finer classification and bounding box regression (second stage). The advantage of this design is that a large number of negative samples are filtered out at the RPN stage, and the second stage can focus on high-quality discrimination on a small number of candidate regions, thus often having an edge in accuracy and being easier to extend to tasks like instance segmentation (Mask R-CNN) and keypoint detection (Keypoint R-CNN). However, the multi-stage structure brings relatively higher computational and implementation complexity, making it more suitable for offline or near-real-time scenarios where real-time requirements are less stringent but accuracy and extensibility are emphasized.Two-stage detection is represented by Faster R-CNN. It first generates a batch of candidate boxes with "high probability of containing an object" through an RPN (Region Proposal Network) on the backbone feature maps (first stage), then performs RoI alignment and feature extraction for each candidate region, followed by finer classification and bounding box regression (second stage). The advantage of this design is that a large number of negative samples are filtered out at the RPN stage, and the second stage can focus on high-quality discrimination on a small number of candidate regions, thus often having an edge in accuracy and being easier to extend to tasks like instance segmentation (Mask R-CNN) and keypoint detection (Keypoint R-CNN). However, the multi-stage structure brings relatively higher computational and implementation complexity, making it more suitable for offline or near-real-time scenarios where real-time requirements are less stringent but accuracy and extensibility are emphasized.
One-stage detection aims to streamline the entire pipeline, completing both category classification and bounding box regression in a unified network. Representative models include SSD, RetinaNet, and the YOLO series: they directly predict "foreground/background + category + bbox" for several candidate boxes at each position on multi-scale feature maps, omitting the explicit proposal stage and being more suitable for end-to-end acceleration and deployment. Early one-stage detectors had a certain accuracy gap compared to two-stage ones, but with their simple structure and fast speed, they quickly dominated industry. With the introduction of FPN, focal loss, IoU-aware loss, and stronger backbones and necks, newer models like RetinaNet, YOLOX, YOLOv7/8/10 have achieved accuracy–speed balances that "approach or even surpass two-stage" on many tasks.One-stage detection aims to streamline the entire pipeline, completing both category classification and bounding box regression in a unified network. Representative models include SSD, RetinaNet, and the YOLO series: they directly predict "foreground/background + category + bbox" for several candidate boxes at each position on multi-scale feature maps, omitting the explicit proposal stage and being more suitable for end-to-end acceleration and deployment. Early one-stage detectors had a certain accuracy gap compared to two-stage ones, but with their simple structure and fast speed, they quickly dominated industry. With the introduction of FPN, focal loss, IoU-aware loss, and stronger backbones and necks, newer models like RetinaNet, YOLOX, YOLOv7/8/10 have achieved accuracy–speed balances that "approach or even surpass two-stage" on many tasks.
At the application level, engineering choices between these two architectures are typically made based on requirements: for cloud-based batch offline analysis tasks requiring high accuracy and extensibility (e.g., simultaneous detection + segmentation + keypoints), two-stage detection remains a stable and reliable choice; for latency-sensitive scenarios like edge devices, mobile applications, and real-time camera detection, one-stage detectors like the YOLO series are almost the default go-to, often combined with techniques like quantization, pruning, and distillation to further compress models and increase throughput.At the application level, engineering choices between these two architectures are typically made based on requirements: for cloud-based batch offline analysis tasks requiring high accuracy and extensibility (e.g., simultaneous detection + segmentation + keypoints), two-stage detection remains a stable and reliable choice; for latency-sensitive scenarios like edge devices, mobile applications, and real-time camera detection, one-stage detectors like the YOLO series are almost the default go-to, often combined with techniques like quantization, pruning, and distillation to further compress models and increase throughput.
On the question of how to define "candidate boxes," detection methods can be divided into two major categories: anchor-based and anchor-free. Early mainstream methods (e.g., Faster R-CNN, SSD, RetinaNet, YOLOv3/v4/v5) adopted the anchor-based approach: predefining several anchor boxes with different scales and aspect ratios at each position on the feature map, then learning the foreground probability and bbox offset for each anchor. This approach is simple to implement and effective, but requires substantial manual tuning of anchor sizes and ratios, and in small-object or dense-object scenarios, can lead to an enormous number of anchors and extreme positive-negative sample imbalance.On the question of how to define "candidate boxes," detection methods can be divided into two major categories: anchor-based and anchor-free. Early mainstream methods (e.g., Faster R-CNN, SSD, RetinaNet, YOLOv3/v4/v5) adopted the anchor-based approach: predefining several anchor boxes with different scales and aspect ratios at each position on the feature map, then learning the foreground probability and bbox offset for each anchor. This approach is simple to implement and effective, but requires substantial manual tuning of anchor sizes and ratios, and in small-object or dense-object scenarios, can lead to an enormous number of anchors and extreme positive-negative sample imbalance.
Anchor-free methods attempt to break free from dependence on predefined anchors. Represented by FCOS, CenterNet, ATSS, etc., they typically directly predict at each pixel of the feature map "whether this is the center of some object (or belongs to that object)" and the corresponding boundary distances, thus completely avoiding the complexity of preset anchors. The benefits are: simpler model structure, more natural training sample assignment strategies, and better generalization and scalability, especially when facing real-world scenes with large scale variations and complex object shapes. At the same time, anchor-free detectors have also driven more pixel/point-based unified frameworks, making it easier to jointly model detection with keypoints, segmentation, and other tasks.Anchor-free methods attempt to break free from dependence on predefined anchors. Represented by FCOS, CenterNet, ATSS, etc., they typically directly predict at each pixel of the feature map "whether this is the center of some object (or belongs to that object)" and the corresponding boundary distances, thus completely avoiding the complexity of preset anchors. The benefits are: simpler model structure, more natural training sample assignment strategies, and better generalization and scalability, especially when facing real-world scenes with large scale variations and complex object shapes. At the same time, anchor-free detectors have also driven more pixel/point-based unified frameworks, making it easier to jointly model detection with keypoints, segmentation, and other tasks.
Going further, Transformer-based detectors like DETR / Deformable DETR rethink the detection problem from another dimension: instead of densely laying anchors on feature maps, they introduce a fixed number of "object queries" and, through the Transformer's self-attention and cross-attention mechanisms, "generate" a set of object predictions from global features, achieving one-to-one alignment via Hungarian matching. This set prediction approach completely eliminates traditional components like NMS and manual sample assignment — conceptually very clean, but early implementations suffered from slow convergence and being unfriendly to small objects. Subsequent work like Deformable DETR, by introducing deformable attention and multi-scale mechanisms, has significantly improved both convergence speed and performance, gradually gaining more applications in detection and multi-task scenarios.Going further, Transformer-based detectors like DETR / Deformable DETR rethink the detection problem from another dimension: instead of densely laying anchors on feature maps, they introduce a fixed number of "object queries" and, through the Transformer's self-attention and cross-attention mechanisms, "generate" a set of object predictions from global features, achieving one-to-one alignment via Hungarian matching. This set prediction approach completely eliminates traditional components like NMS and manual sample assignment — conceptually very clean, but early implementations suffered from slow convergence and being unfriendly to small objects. Subsequent work like Deformable DETR, by introducing deformable attention and multi-scale mechanisms, has significantly improved both convergence speed and performance, gradually gaining more applications in detection and multi-task scenarios.
For engineering practice, anchor-based, anchor-free, and Transformer detection are not mutually exclusive choices but rather an evolutionary chain: from heavily engineered anchor design, to more end-to-end point/center prediction, to fully set-prediction and attention-based unified frameworks. In current industrial deployment, mature anchor-based models like the YOLO series remain the workhorses, while anchor-free and the DETR family appear more in systems with higher demands for structural simplicity, multi-task unification, and scalability.For engineering practice, anchor-based, anchor-free, and Transformer detection are not mutually exclusive choices but rather an evolutionary chain: from heavily engineered anchor design, to more end-to-end point/center prediction, to fully set-prediction and attention-based unified frameworks. In current industrial deployment, mature anchor-based models like the YOLO series remain the workhorses, while anchor-free and the DETR family appear more in systems with higher demands for structural simplicity, multi-task unification, and scalability.
Object detection on public datasets often gives the impression that "the problem is basically solved," but once entering real-world scenarios, two types of challenging problems immediately arise: small/dense objects and robust detection and tracking in video.Object detection on public datasets often gives the impression that "the problem is basically solved," but once entering real-world scenarios, two types of challenging problems immediately arise: small/dense objects and robust detection and tracking in video.
In small object detection, the target often occupies only a tiny pixel area in the original image — for example, distant pedestrians, faraway vehicles, aerial drones, or minute defects on high-resolution industrial images. As the backbone downsamples and feature map resolution decreases, these small objects can easily be "drowned out" in high-level features, leading to missed detections. To address this, detectors typically employ multi-scale feature pyramids (FPN/PAFPN, etc.), increase input resolution, add detection heads on shallow feature maps, and even design dedicated branches and loss weighting strategies for small objects. At the data level, techniques like cropping, zooming, and small-object resampling are also needed to improve the model's perception and memory of small-scale targets.In small object detection, the target often occupies only a tiny pixel area in the original image — for example, distant pedestrians, faraway vehicles, aerial drones, or minute defects on high-resolution industrial images. As the backbone downsamples and feature map resolution decreases, these small objects can easily be "drowned out" in high-level features, leading to missed detections. To address this, detectors typically employ multi-scale feature pyramids (FPN/PAFPN, etc.), increase input resolution, add detection heads on shallow feature maps, and even design dedicated branches and loss weighting strategies for small objects. At the data level, techniques like cropping, zooming, and small-object resampling are also needed to improve the model's perception and memory of small-scale targets.
Dense objects (e.g., crowded crowds, dense parking lots, tightly packed products/parts) expose problems like anchor box overlap, NMS false suppression, and severe occlusion. Improvement strategies include finer label assignment (e.g., adaptive assignment methods like ATSS), soft NMS or learning-based deduplication strategies, and mitigating inter-box competition through center-point/density-map modeling. In industrial quality inspection, many systems also combine detection with pixel-level segmentation for more precise defect localization, facilitating subsequent automated processing.Dense objects (e.g., crowded crowds, dense parking lots, tightly packed products/parts) expose problems like anchor box overlap, NMS false suppression, and severe occlusion. Improvement strategies include finer label assignment (e.g., adaptive assignment methods like ATSS), soft NMS or learning-based deduplication strategies, and mitigating inter-box competition through center-point/density-map modeling. In industrial quality inspection, many systems also combine detection with pixel-level segmentation for more precise defect localization, facilitating subsequent automated processing.
When detection extends from single frames to video, another challenge is temporal continuity and target stability. Single-frame detectors make independent predictions on each frame, making it difficult to avoid short-term missed detections, ID jitter, and false alarms — yet real-world applications for alerting, counting, and trajectory analysis often require consistent cross-frame target trajectories. For this, video object detection typically overlays a tracking module, connecting "detection + object tracking": the classic approach uses an image detector as the frontend, and in the backend uses Kalman filtering, Hungarian matching, and appearance feature similarity for multi-object tracking (e.g., SORT, DeepSORT). More advanced approaches integrate tracking heads directly into the detection network, jointly learning detection and cross-frame association to improve robustness in scenarios with short-term occlusion and fast motion.When detection extends from single frames to video, another challenge is temporal continuity and target stability. Single-frame detectors make independent predictions on each frame, making it difficult to avoid short-term missed detections, ID jitter, and false alarms — yet real-world applications for alerting, counting, and trajectory analysis often require consistent cross-frame target trajectories. For this, video object detection typically overlays a tracking module, connecting "detection + object tracking": the classic approach uses an image detector as the frontend, and in the backend uses Kalman filtering, Hungarian matching, and appearance feature similarity for multi-object tracking (e.g., SORT, DeepSORT). More advanced approaches integrate tracking heads directly into the detection network, jointly learning detection and cross-frame association to improve robustness in scenarios with short-term occlusion and fast motion.
In real systems, small objects, dense objects, and video detection are often not isolated problems but appear simultaneously: for example, distant pedestrians/vehicles in urban road surveillance, dense crowds in station squares, and high-speed moving parts in production line video. This also means that high-quality object detection modules, beyond having impressive metrics on standard benchmarks, need to withstand the test of various complex factors under real-world conditions like multi-scale, multi-density, and long-duration video, in order to truly support upper-level behavior analysis, intelligent alerting, and multimodal understanding.In real systems, small objects, dense objects, and video detection are often not isolated problems but appear simultaneously: for example, distant pedestrians/vehicles in urban road surveillance, dense crowds in station squares, and high-speed moving parts in production line video. This also means that high-quality object detection modules, beyond having impressive metrics on standard benchmarks, need to withstand the test of various complex factors under real-world conditions like multi-scale, multi-density, and long-duration video, in order to truly support upper-level behavior analysis, intelligent alerting, and multimodal understanding.
With object detection, we already know "what objects are in the image and roughly where they are," but many tasks require even finer structured understanding: precisely determining, for every pixel, which category it belongs to and which instance it belongs to. For example, in autonomous driving, we need to know which pixels are road, which are people, and which are cars; cutout tools need to cleanly separate hair strands from the background; medical images require precise delineation of tumor and organ boundaries. These tasks are collectively called image segmentation, which directly outputs semantic or instance labels at the pixel level, providing finer-grained spatial structure information compared to detection.With object detection, we already know "what objects are in the image and roughly where they are," but many tasks require even finer structured understanding: precisely determining, for every pixel, which category it belongs to and which instance it belongs to. For example, in autonomous driving, we need to know which pixels are road, which are people, and which are cars; cutout tools need to cleanly separate hair strands from the background; medical images require precise delineation of tumor and organ boundaries. These tasks are collectively called image segmentation, which directly outputs semantic or instance labels at the pixel level, providing finer-grained spatial structure information compared to detection.
From a product perspective, image segmentation is the core capability for "pixel-level structuring": cutout and background replacement tools rely on it to decide which pixels to keep; autonomous driving perception modules rely on it to build fine-grained "drivable area + obstacle" maps; medical imaging software relies on it to measure lesion size, shape, and volume; remote sensing platforms rely on it to distinguish farmland, water bodies, buildings, roads, and other land features. Below, we organize image segmentation from three angles — scenarios, principles, and models — and expand on directions like semantic/instance/panoptic/large-model segmentation in subsequent subsections.From a product perspective, image segmentation is the core capability for "pixel-level structuring": cutout and background replacement tools rely on it to decide which pixels to keep; autonomous driving perception modules rely on it to build fine-grained "drivable area + obstacle" maps; medical imaging software relies on it to measure lesion size, shape, and volume; remote sensing platforms rely on it to distinguish farmland, water bodies, buildings, roads, and other land features. Below, we organize image segmentation from three angles — scenarios, principles, and models — and expand on directions like semantic/instance/panoptic/large-model segmentation in subsequent subsections.
Image segmentation is essentially "dense prediction": extracting multi-scale features from the input image through an encoder (backbone), then progressively restoring the feature maps to a segmentation map of the same size as the input through a decoder or upsampling modules, outputting a semantic or instance label at each pixel position.Image segmentation is essentially "dense prediction": extracting multi-scale features from the input image through an encoder (backbone), then progressively restoring the feature maps to a segmentation map of the same size as the input through a decoder or upsampling modules, outputting a semantic or instance label at each pixel position.
Compared to detection, segmentation is more sensitive to spatial details and boundary quality, requiring richer multi-scale contextual information and finer upsampling/fusion strategies.Compared to detection, segmentation is more sensitive to spatial details and boundary quality, requiring richer multi-scale contextual information and finer upsampling/fusion strategies.
Classic to latest segmentation models have roughly evolved along the path of "FCN → encoder–decoder → multi-scale context → detection+segmentation unification → large-model segmentation":Classic to latest segmentation models have roughly evolved along the path of "FCN → encoder–decoder → multi-scale context → detection+segmentation unification → large-model segmentation":
Overall, image segmentation provides finer spatial structure representation compared to object detection, and is an indispensable link in building highly reliable perception systems and advanced editing tools. Below, we expand on three directions: semantic segmentation and instance segmentation, panoptic segmentation and detection unification, and general segmentation, large models, and unsupervised segmentation.Overall, image segmentation provides finer spatial structure representation compared to object detection, and is an indispensable link in building highly reliable perception systems and advanced editing tools. Below, we expand on three directions: semantic segmentation and instance segmentation, panoptic segmentation and detection unification, and general segmentation, large models, and unsupervised segmentation.
The goal of semantic segmentation is to assign a semantic category to every pixel in the image, so that the network learns "this region is road, that region is car, here is person, there is sky and building." Classic approaches typically use an encoder–decoder structure: the encoder (e.g., ResNet, EfficientNet, Swin Transformer) extracts progressively downsampled high-level features, and the decoder restores resolution to the original through upsampling, skip connections, and multi-scale fusion, combining coarse high-level semantic features with low-level details. FCN was the first to systematize this dense prediction form; U-Net achieved great success in medical imaging through its symmetric U-shaped structure and abundant skip connections; the DeepLab series expanded the receptive field without reducing resolution through dilated convolutions and ASPP (Atrous Spatial Pyramid Pooling); PSPNet obtained global context information through pyramid pooling. These models collectively drove large-scale applications in road scenes, remote sensing, medical imaging, and other domains.The goal of semantic segmentation is to assign a semantic category to every pixel in the image, so that the network learns "this region is road, that region is car, here is person, there is sky and building." Classic approaches typically use an encoder–decoder structure: the encoder (e.g., ResNet, EfficientNet, Swin Transformer) extracts progressively downsampled high-level features, and the decoder restores resolution to the original through upsampling, skip connections, and multi-scale fusion, combining coarse high-level semantic features with low-level details. FCN was the first to systematize this dense prediction form; U-Net achieved great success in medical imaging through its symmetric U-shaped structure and abundant skip connections; the DeepLab series expanded the receptive field without reducing resolution through dilated convolutions and ASPP (Atrous Spatial Pyramid Pooling); PSPNet obtained global context information through pyramid pooling. These models collectively drove large-scale applications in road scenes, remote sensing, medical imaging, and other domains.
Instance segmentation further distinguishes different individuals of the same class on top of pixel semantic labels: not only knowing which pixels are "car," but also knowing which pixels belong to which specific car. The most representative model is Mask R-CNN, which adds a parallel segmentation branch to the Faster R-CNN detection framework: the detection head first predicts the category and position of each candidate box, then generates a binary mask within each box, yielding "box + mask" object-level segmentation results. Compared to pure semantic segmentation, this approach handles object overlap and occlusion well, and serves as the foundation for tasks like portrait/product cutout, multi-object counting, and fine-grained editing. Subsequent instance segmentation methods have continuously improved in mask quality, multi-scale handling, and speed, with new anchor-free and Transformer-based architectures also emerging, but the "detection + local segmentation" approach remains very mainstream.Instance segmentation further distinguishes different individuals of the same class on top of pixel semantic labels: not only knowing which pixels are "car," but also knowing which pixels belong to which specific car. The most representative model is Mask R-CNN, which adds a parallel segmentation branch to the Faster R-CNN detection framework: the detection head first predicts the category and position of each candidate box, then generates a binary mask within each box, yielding "box + mask" object-level segmentation results. Compared to pure semantic segmentation, this approach handles object overlap and occlusion well, and serves as the foundation for tasks like portrait/product cutout, multi-object counting, and fine-grained editing. Subsequent instance segmentation methods have continuously improved in mask quality, multi-scale handling, and speed, with new anchor-free and Transformer-based architectures also emerging, but the "detection + local segmentation" approach remains very mainstream.
At the product level, semantic segmentation typically appears in "scene-level" applications such as autonomous driving road segmentation, remote sensing land feature recognition, and medical organ segmentation; instance segmentation is more commonly used for "object-level" cutout, counting, and editing, such as one-click selection and separation of each car, each person, and each product. The combination of both provides fine-grained yet structured spatial information for upper-level tasks.At the product level, semantic segmentation typically appears in "scene-level" applications such as autonomous driving road segmentation, remote sensing land feature recognition, and medical organ segmentation; instance segmentation is more commonly used for "object-level" cutout, counting, and editing, such as one-click selection and separation of each car, each person, and each product. The combination of both provides fine-grained yet structured spatial information for upper-level tasks.
Doing only semantic segmentation mixes same-class objects together (all "car" pixels belong to the same class); doing only instance segmentation often focuses only on countable "things" (e.g., people, cars, animals) while ignoring large areas of uncountable "stuff" (e.g., road, grass, sky). In many scenarios, we need both instance-level masks for each object and an understanding of the overall scene composition. This gave rise to panoptic segmentation: simultaneously providing a semantic class and instance ID for every pixel, achieving unified modeling of things + stuff.Doing only semantic segmentation mixes same-class objects together (all "car" pixels belong to the same class); doing only instance segmentation often focuses only on countable "things" (e.g., people, cars, animals) while ignoring large areas of uncountable "stuff" (e.g., road, grass, sky). In many scenarios, we need both instance-level masks for each object and an understanding of the overall scene composition. This gave rise to panoptic segmentation: simultaneously providing a semantic class and instance ID for every pixel, achieving unified modeling of things + stuff.
Early panoptic segmentation systems were typically implemented through "semantic segmentation model + instance segmentation model + post-processing fusion": first using one network to predict the semantic category of each pixel, then using another network to output masks and categories for each instance, and finally merging the two into a consistent panoptic segmentation result through a set of rules (e.g., priority, overlap handling). Panoptic FPN represents a more elegant engineering path: on a shared backbone and feature pyramid (FPN), separately attaching semantic segmentation heads and instance segmentation heads, obtaining both outputs simultaneously through joint training and feature sharing, then fusing them through lightweight post-processing. This not only improves efficiency but also enhances consistency between semantics and instances.Early panoptic segmentation systems were typically implemented through "semantic segmentation model + instance segmentation model + post-processing fusion": first using one network to predict the semantic category of each pixel, then using another network to output masks and categories for each instance, and finally merging the two into a consistent panoptic segmentation result through a set of rules (e.g., priority, overlap handling). Panoptic FPN represents a more elegant engineering path: on a shared backbone and feature pyramid (FPN), separately attaching semantic segmentation heads and instance segmentation heads, obtaining both outputs simultaneously through joint training and feature sharing, then fusing them through lightweight post-processing. This not only improves efficiency but also enhances consistency between semantics and instances.
At the model level, with the development of detection/segmentation unification and Transformer architectures, unified panoptic segmentation frameworks like Mask2Former have emerged: they tend to use a universal "query + mask decoder" structure, simultaneously predicting masks for semantics, instances, and even other downstream tasks within the same network, thereby greatly simplifying the system architecture and facilitating multi-task extension. For complex tasks like autonomous driving, robot navigation, and AR scene understanding, panoptic segmentation provides a more complete scene description closer to "human subjective perception," allowing upper-level decision-making and planning to operate on more accurate spatial semantics.At the model level, with the development of detection/segmentation unification and Transformer architectures, unified panoptic segmentation frameworks like Mask2Former have emerged: they tend to use a universal "query + mask decoder" structure, simultaneously predicting masks for semantics, instances, and even other downstream tasks within the same network, thereby greatly simplifying the system architecture and facilitating multi-task extension. For complex tasks like autonomous driving, robot navigation, and AR scene understanding, panoptic segmentation provides a more complete scene description closer to "human subjective perception," allowing upper-level decision-making and planning to operate on more accurate spatial semantics.
In product form, panoptic segmentation is often embedded within autonomous driving, robotics systems, and high-end visual analysis platforms. Users may not directly perceive the concept of "panoptic segmentation," but they genuinely benefit from more robust scene understanding and more natural interactive experiences.In product form, panoptic segmentation is often embedded within autonomous driving, robotics systems, and high-end visual analysis platforms. Users may not directly perceive the concept of "panoptic segmentation," but they genuinely benefit from more robust scene understanding and more natural interactive experiences.
Traditional segmentation models are often trained around specific datasets and tasks: for example, "19-class semantic segmentation of road scenes," "segmentation of a certain type of tumor," "segmentation of certain product categories," etc. Each new task requires re-annotation and re-training. In real business, this approach of heavily relying on precisely labeled data is extremely costly and struggles to cover long-tail categories and constantly emerging new scenarios. In recent years, with the development of large-scale pre-trained vision models and prompt-based paradigms, general segmentation large models represented by the Segment Anything Model (SAM) have emerged, attempting to elevate segmentation capability from "task-specific customization" to "infrastructure."Traditional segmentation models are often trained around specific datasets and tasks: for example, "19-class semantic segmentation of road scenes," "segmentation of a certain type of tumor," "segmentation of certain product categories," etc. Each new task requires re-annotation and re-training. In real business, this approach of heavily relying on precisely labeled data is extremely costly and struggles to cover long-tail categories and constantly emerging new scenarios. In recent years, with the development of large-scale pre-trained vision models and prompt-based paradigms, general segmentation large models represented by the Segment Anything Model (SAM) have emerged, attempting to elevate segmentation capability from "task-specific customization" to "infrastructure."
Taking SAM as an example, it learns universal features of the entire image through a powerful image encoder (typically a large-scale pre-trained ViT), and then converts user-provided points, boxes, text prompts, etc. into segmentation results through a lightweight prompt encoder and mask decoder. During training, SAM leverages massive, multi-source, multi-task mask annotations, so that the model learns a "generalized segmentation capability" rather than rote memorization of a specific dataset's labels. During use, users only need to provide very few prompts (a single point or a rough box) to obtain high-quality masks on various unseen image types and object categories. This paradigm greatly lowers the barrier to building new segmentation applications and provides a powerful tool for unsupervised/weakly-supervised scenarios.Taking SAM as an example, it learns universal features of the entire image through a powerful image encoder (typically a large-scale pre-trained ViT), and then converts user-provided points, boxes, text prompts, etc. into segmentation results through a lightweight prompt encoder and mask decoder. During training, SAM leverages massive, multi-source, multi-task mask annotations, so that the model learns a "generalized segmentation capability" rather than rote memorization of a specific dataset's labels. During use, users only need to provide very few prompts (a single point or a rough box) to obtain high-quality masks on various unseen image types and object categories. This paradigm greatly lowers the barrier to building new segmentation applications and provides a powerful tool for unsupervised/weakly-supervised scenarios.
Related to this is the broader direction of unsupervised / self-supervised segmentation: without relying on or minimally relying on human-annotated masks, automatically dividing images into several meaningful regions through signals such as intra-image similarity, temporal consistency, and multi-view constraints. Early work focused more on "visual clustering" and region proposal generation, while nowadays it is increasingly internalized by large models as a form of representation learning, providing good initialization for downstream segmentation tasks. Combined with text–image contrastive learning models like CLIP, an increasing number of methods can perform zero-shot or few-shot segmentation under the condition of "only providing text category names, without providing mask annotations," offering new solutions for cold-start scenarios and long-tail categories.Related to this is the broader direction of unsupervised / self-supervised segmentation: without relying on or minimally relying on human-annotated masks, automatically dividing images into several meaningful regions through signals such as intra-image similarity, temporal consistency, and multi-view constraints. Early work focused more on "visual clustering" and region proposal generation, while nowadays it is increasingly internalized by large models as a form of representation learning, providing good initialization for downstream segmentation tasks. Combined with text–image contrastive learning models like CLIP, an increasing number of methods can perform zero-shot or few-shot segmentation under the condition of "only providing text category names, without providing mask annotations," offering new solutions for cold-start scenarios and long-tail categories.
In actual products, general segmentation large models often appear in forms like "interactive cutout tools," "smart selection," and "one-click background removal," and are gradually being integrated into professional software in fields like medicine, remote sensing, and industry, serving as accelerators for semi-automatic annotation and assisted segmentation. Compared to traditional custom models, they may not achieve the ultimate performance on a specific task, but they have a significant advantage in "being able to do a bit of everything and quickly landing in multiple scenarios," laying the groundwork for building truly multimodal foundation vision models.In actual products, general segmentation large models often appear in forms like "interactive cutout tools," "smart selection," and "one-click background removal," and are gradually being integrated into professional software in fields like medicine, remote sensing, and industry, serving as accelerators for semi-automatic annotation and assisted segmentation. Compared to traditional custom models, they may not achieve the ultimate performance on a specific task, but they have a significant advantage in "being able to do a bit of everything and quickly landing in multiple scenarios," laying the groundwork for building truly multimodal foundation vision models.
After classification, detection, and segmentation, we already know "what is in the image, where it is, and what each pixel belongs to." But in many real-world tasks, businesses care not only about "object presence and position" but also about pose and action: Is a person walking or running? Is this hand raised or making a certain gesture? Is the worker correctly wearing safety equipment and performing standardized actions? Is the athlete's technique correct? These questions require us to further understand the internal structure and temporal changes of objects.After classification, detection, and segmentation, we already know "what is in the image, where it is, and what each pixel belongs to." But in many real-world tasks, businesses care not only about "object presence and position" but also about pose and action: Is a person walking or running? Is this hand raised or making a certain gesture? Is the worker correctly wearing safety equipment and performing standardized actions? Is the athlete's technique correct? These questions require us to further understand the internal structure and temporal changes of objects.
Keypoint detection and action recognition are two layers of capability designed for this need:Keypoint detection and action recognition are two layers of capability designed for this need:
From a product perspective, this capability widely serves: human-computer interaction (gesture control), sports analysis (technique evaluation), security (fall detection, fighting/running and other abnormal behavior recognition), industrial safety (violation action detection), virtual human driving (driving 3D skeletons and animations via body/facial keypoints), and other scenarios. Below, we organize this layer of capability from three angles — scenarios, principles, and models — and expand on keypoint detection and action recognition in subsections.From a product perspective, this capability widely serves: human-computer interaction (gesture control), sports analysis (technique evaluation), security (fall detection, fighting/running and other abnormal behavior recognition), industrial safety (violation action detection), virtual human driving (driving 3D skeletons and animations via body/facial keypoints), and other scenarios. Below, we organize this layer of capability from three angles — scenarios, principles, and models — and expand on keypoint detection and action recognition in subsections.
The two types of tasks respectively emphasize spatial structure and temporal change, but both are essentially structured predictions in high-dimensional feature space:The two types of tasks respectively emphasize spatial structure and temporal change, but both are essentially structured predictions in high-dimensional feature space:
Common models have roughly developed along the unified paradigm of "convolutional/Transformer feature extraction + keypoint/temporal heads":Common models have roughly developed along the unified paradigm of "convolutional/Transformer feature extraction + keypoint/temporal heads":
Below, we expand on two directions: keypoint detection and pose estimation and action recognition and behavior understanding.Below, we expand on two directions: keypoint detection and pose estimation and action recognition and behavior understanding.
Keypoint detection (also often called pose estimation) focuses on the spatial structure in a single frame or single image: finding a set of semantically meaningful keypoints in a 2D image and connecting them into a skeleton. For example, in human pose estimation, we typically need to detect joints like the head, shoulders, elbows, wrists, hips, knees, and ankles; in facial pose, it's the eye corners, mouth corners, nose tip, face contour, etc.; in hand pose, it's the finger roots, knuckles, and fingertips. For non-human objects like robotic arms and articulated structural components, a keypoint system can also be defined similarly.Keypoint detection (also often called pose estimation) focuses on the spatial structure in a single frame or single image: finding a set of semantically meaningful keypoints in a 2D image and connecting them into a skeleton. For example, in human pose estimation, we typically need to detect joints like the head, shoulders, elbows, wrists, hips, knees, and ankles; in facial pose, it's the eye corners, mouth corners, nose tip, face contour, etc.; in hand pose, it's the finger roots, knuckles, and fingertips. For non-human objects like robotic arms and articulated structural components, a keypoint system can also be defined similarly.
In model design, keypoint detection commonly uses the "feature extraction + heatmap prediction" paradigm:In model design, keypoint detection commonly uses the "feature extraction + heatmap prediction" paradigm:
For multi-person scenarios, pose estimation methods roughly divide into two paths:For multi-person scenarios, pose estimation methods roughly divide into two paths:
In recent years, Transformer-based pose estimation models have also gradually emerged, treating keypoint detection as a set of "query–response" tasks, similar to DETR, which can unify object detection and pose estimation architecturally. In engineering applications, keypoint detection capability is typically packaged as "human body/gesture/facial keypoint SDK or API," where upstream applications only need to input images or video frames to obtain structured skeleton coordinates for subsequent action recognition, interaction control, or animation driving.In recent years, Transformer-based pose estimation models have also gradually emerged, treating keypoint detection as a set of "query–response" tasks, similar to DETR, which can unify object detection and pose estimation architecturally. In engineering applications, keypoint detection capability is typically packaged as "human body/gesture/facial keypoint SDK or API," where upstream applications only need to input images or video frames to obtain structured skeleton coordinates for subsequent action recognition, interaction control, or animation driving.
After obtaining keypoints or high-level visual features, the next step is to understand changes in the temporal dimension — that is, action recognition and behavior analysis. Unlike keypoint detection, action recognition is no longer confined to single frames; it is concerned with the evolution patterns of features over a period of time: from "raising a hand" to "waving," from "walking" to "running," from "standing" to "falling."After obtaining keypoints or high-level visual features, the next step is to understand changes in the temporal dimension — that is, action recognition and behavior analysis. Unlike keypoint detection, action recognition is no longer confined to single frames; it is concerned with the evolution patterns of features over a period of time: from "raising a hand" to "waving," from "walking" to "running," from "standing" to "falling."
In terms of input representation, there are roughly three routes:In terms of input representation, there are roughly three routes:
Correspondingly, model structures have also shown diversified development:Correspondingly, model structures have also shown diversified development:
On the business side, action recognition is often combined with detection, tracking, and keypoint detection to form end-to-end behavior analysis systems:On the business side, action recognition is often combined with detection, tracking, and keypoint detection to form end-to-end behavior analysis systems:
Looking ahead, multimodal large models are elevating "action recognition" to a higher level of "event and intent understanding": models can not only label "walking, running, making a phone call" but also answer descriptions closer to everyday language like "this person seems to be signaling someone" or "these two people are having an argument." Keypoint detection and action recognition, as important structured motion cues, together with appearance features and language prompts, jointly support more complex spatiotemporal understanding capabilities.Looking ahead, multimodal large models are elevating "action recognition" to a higher level of "event and intent understanding": models can not only label "walking, running, making a phone call" but also answer descriptions closer to everyday language like "this person seems to be signaling someone" or "these two people are having an argument." Keypoint detection and action recognition, as important structured motion cues, together with appearance features and language prompts, jointly support more complex spatiotemporal understanding capabilities.
The detection and segmentation capabilities discussed earlier all basically default to one premise: the set of categories during training and inference is fixed. That is, the model has fully seen "all categories to be recognized" during the training phase, and during inference, it only needs to choose from this closed set of labels. But the real world is far more complex than datasets: new products, new brands, new signs, new species, and new scenarios appear at any time, and it is impossible to prepare sufficient annotated data for each new class and retrain the detector. This gave rise to open-vocabulary / open-world / open-domain detection: under the condition that training data only covers a limited set of "known classes," enabling the model to still perceive, locate, and recognize unseen new classes during inference, while maintaining robustness under changes in visual style and capture domain.The detection and segmentation capabilities discussed earlier all basically default to one premise: the set of categories during training and inference is fixed. That is, the model has fully seen "all categories to be recognized" during the training phase, and during inference, it only needs to choose from this closed set of labels. But the real world is far more complex than datasets: new products, new brands, new signs, new species, and new scenarios appear at any time, and it is impossible to prepare sufficient annotated data for each new class and retrain the detector. This gave rise to open-vocabulary / open-world / open-domain detection: under the condition that training data only covers a limited set of "known classes," enabling the model to still perceive, locate, and recognize unseen new classes during inference, while maintaining robustness under changes in visual style and capture domain.
You can think of this layer as: adding "alignment and generalization capability to the language space and open world" on top of traditional detection. The model no longer only says "this is one of the 80 COCO classes," but can understand and retrieve objects in the space of arbitrary text descriptions — for example, "detect all 'red sports shoes' in the image," "mark all 'suspected small aircraft'," even if these fine-grained categories never explicitly appeared in the training set. Below, we organize this layer from three angles — scenarios, principles, and models — and expand on open-vocabulary detection, open-world detection, and open-domain generalization in subsections.You can think of this layer as: adding "alignment and generalization capability to the language space and open world" on top of traditional detection. The model no longer only says "this is one of the 80 COCO classes," but can understand and retrieve objects in the space of arbitrary text descriptions — for example, "detect all 'red sports shoes' in the image," "mark all 'suspected small aircraft'," even if these fine-grained categories never explicitly appeared in the training set. Below, we organize this layer from three angles — scenarios, principles, and models — and expand on open-vocabulary detection, open-world detection, and open-domain generalization in subsections.
The core of these methods is to replace the traditional "fixed one-hot category head" with a vision–language aligned embedding space, and handle "unseen classes" and "new domains" through various mechanisms:The core of these methods is to replace the traditional "fixed one-hot category head" with a vision–language aligned embedding space, and handle "unseen classes" and "new domains" through various mechanisms:
Current mainstream technical routes for open-vocabulary / open-world / open-domain detection basically revolve around "large-scale vision–language pre-training + detection head adaptation + domain generalization mechanisms":Current mainstream technical routes for open-vocabulary / open-world / open-domain detection basically revolve around "large-scale vision–language pre-training + detection head adaptation + domain generalization mechanisms":
In specific product forms, open-vocabulary/open-world/open-domain detection often manifests as "more natural, less restricted" visual interfaces: users don't need to pre-agree on a small set of fixed labels, but can describe what they want to find in natural language; the system also doesn't need to retrain detectors from scratch for each business scenario, but can quickly adapt through prompts or few-shot samples based on a unified general model. For large-scale product/species recognition and globally deployed security and autonomous driving perception systems, this layer of capability is becoming the key springboard from "closed dataset performance" to "real open-world usability."In specific product forms, open-vocabulary/open-world/open-domain detection often manifests as "more natural, less restricted" visual interfaces: users don't need to pre-agree on a small set of fixed labels, but can describe what they want to find in natural language; the system also doesn't need to retrain detectors from scratch for each business scenario, but can quickly adapt through prompts or few-shot samples based on a unified general model. For large-scale product/species recognition and globally deployed security and autonomous driving perception systems, this layer of capability is becoming the key springboard from "closed dataset performance" to "real open-world usability."
The starting point of open-vocabulary detection is to break through the limitation of "fixed category heads" in traditional detection. Previous detectors attached a fixed-size classification layer at the top (corresponding to the N categories in the training set), and after training, could only choose among these N categories. Open-vocabulary detection, by introducing text encoders and a shared semantic embedding space, allows the region features output by the detection head to be compared for similarity with arbitrary text descriptions, thereby accepting unseen new categories during inference.The starting point of open-vocabulary detection is to break through the limitation of "fixed category heads" in traditional detection. Previous detectors attached a fixed-size classification layer at the top (corresponding to the N categories in the training set), and after training, could only choose among these N categories. Open-vocabulary detection, by introducing text encoders and a shared semantic embedding space, allows the region features output by the detection head to be compared for similarity with arbitrary text descriptions, thereby accepting unseen new categories during inference.
A typical approach uses CLIP-like vision–language pre-trained models:A typical approach uses CLIP-like vision–language pre-trained models:
During inference, the system no longer relies on a fixed set of class names from training, but allows users to provide arbitrary category words or natural language descriptions online, convert them to embeddings via the text encoder, and then perform similarity matching with region features. This enables the detector to support flexible requirements like "detect all skateboards," "detect all green plants," "detect all safety-related equipment" without retraining — even if certain specific categories never had complete annotations in the training set, as long as they semantically overlap with the pre-trained image–text space, they can be recognized and located to some degree.During inference, the system no longer relies on a fixed set of class names from training, but allows users to provide arbitrary category words or natural language descriptions online, convert them to embeddings via the text encoder, and then perform similarity matching with region features. This enables the detector to support flexible requirements like "detect all skateboards," "detect all green plants," "detect all safety-related equipment" without retraining — even if certain specific categories never had complete annotations in the training set, as long as they semantically overlap with the pre-trained image–text space, they can be recognized and located to some degree.
In engineering practice, open-vocabulary detection needs to balance effectiveness and efficiency: on one hand, maintaining semantic alignment with the large-scale pre-trained vision–language backbone; on the other hand, meeting the detection task's requirements for multi-scale and real-time performance. Mainstream CLIP-based detectors often adopt the approach of "pre-computed text embeddings + efficient vector similarity computation" to avoid repeatedly encoding text during online serving, while also quantizing or distilling region features to balance accuracy and inference speed.In engineering practice, open-vocabulary detection needs to balance effectiveness and efficiency: on one hand, maintaining semantic alignment with the large-scale pre-trained vision–language backbone; on the other hand, meeting the detection task's requirements for multi-scale and real-time performance. Mainstream CLIP-based detectors often adopt the approach of "pre-computed text embeddings + efficient vector similarity computation" to avoid repeatedly encoding text during online serving, while also quantizing or distilling region features to balance accuracy and inference speed.
Open-world detection, building on open-vocabulary, further requires the model to explicitly handle "unknown classes": the training data only annotates some categories, while other objects are either unannotated or collectively labeled as background. During inference, these "unannotated real objects" should neither be simply treated as background nor incorrectly classified into known categories, but should be detected as "unknown" and have the potential to be subsequently converted into "new known classes."Open-world detection, building on open-vocabulary, further requires the model to explicitly handle "unknown classes": the training data only annotates some categories, while other objects are either unannotated or collectively labeled as background. During inference, these "unannotated real objects" should neither be simply treated as background nor incorrectly classified into known categories, but should be detected as "unknown" and have the potential to be subsequently converted into "new known classes."
In modeling, open-world detection typically needs to address three problems:In modeling, open-world detection typically needs to address three problems:
From a product perspective, open-world detection is particularly suitable for scenarios where categories are constantly growing and long-tail is extremely severe, such as natural species recognition, product recognition for rapidly emerging new items, and abnormal object detection in complex security scenarios. The system can first use open-world detection to mark "any non-background suspicious objects," and gradually upgrade valuable clusters to formal categories through manual or semi-automatic annotation, thus forming a detection system where "categories can sustainably grow," rather than being constrained by fixed datasets.From a product perspective, open-world detection is particularly suitable for scenarios where categories are constantly growing and long-tail is extremely severe, such as natural species recognition, product recognition for rapidly emerging new items, and abnormal object detection in complex security scenarios. The system can first use open-world detection to mark "any non-background suspicious objects," and gradually upgrade valuable clusters to formal categories through manual or semi-automatic annotation, thus forming a detection system where "categories can sustainably grow," rather than being constrained by fixed datasets.
Even if the category set remains unchanged, detectors still encounter severe domain shift in real deployment: training data may come from high-definition daytime cameras in a few cities, while deployment environments include different countries, rural areas, highways, tunnels, nighttime, rain/snow, low-resolution cameras, fisheye lenses, and even infrared imaging. Huge differences also exist between e-commerce product photography and user-taken photos, or between ad images/illustrations/anime styles. Open-domain detection focuses precisely on: maintaining stable and reliable detection performance under conditions of significantly changed image distributions.Even if the category set remains unchanged, detectors still encounter severe domain shift in real deployment: training data may come from high-definition daytime cameras in a few cities, while deployment environments include different countries, rural areas, highways, tunnels, nighttime, rain/snow, low-resolution cameras, fisheye lenses, and even infrared imaging. Huge differences also exist between e-commerce product photography and user-taken photos, or between ad images/illustrations/anime styles. Open-domain detection focuses precisely on: maintaining stable and reliable detection performance under conditions of significantly changed image distributions.
Typical technical paths include:Typical technical paths include:
These open-domain mechanisms are often superimposed with open-vocabulary/open-world capabilities: a general detection system for the real world needs to both understand users' natural language category descriptions (open-vocabulary), give reasonable "unknown" judgments and progressive absorption for newly appearing objects (open-world), and maintain performance across different countries, devices, weather conditions, and styles (open-domain). In engineering deployment, these three are not isolated research directions but together form the key capability combination for moving from "closed benchmarks" to "open-world usability."These open-domain mechanisms are often superimposed with open-vocabulary/open-world capabilities: a general detection system for the real world needs to both understand users' natural language category descriptions (open-vocabulary), give reasonable "unknown" judgments and progressive absorption for newly appearing objects (open-world), and maintain performance across different countries, devices, weather conditions, and styles (open-domain). In engineering deployment, these three are not isolated research directions but together form the key capability combination for moving from "closed benchmarks" to "open-world usability."
The previous sections mainly focused on "single-modality vision": the input is an image, and the output is detection boxes, segmentation masks, category labels, or quality scores. But in many real applications, visual information does not exist in isolation — an image often comes with captions, descriptive text, dialogue, or search queries. Users want to ask "what is happening in this image" or "does this image match this sentence." Vision–language tasks are precisely designed to solve these problems: they take image + text as input or output, and through cross-modal alignment and joint modeling, enable the system to "describe images in words," "answer questions about images," and "find images by text / find text by image."The previous sections mainly focused on "single-modality vision": the input is an image, and the output is detection boxes, segmentation masks, category labels, or quality scores. But in many real applications, visual information does not exist in isolation — an image often comes with captions, descriptive text, dialogue, or search queries. Users want to ask "what is happening in this image" or "does this image match this sentence." Vision–language tasks are precisely designed to solve these problems: they take image + text as input or output, and through cross-modal alignment and joint modeling, enable the system to "describe images in words," "answer questions about images," and "find images by text / find text by image."
From a product perspective, vision–language models (VLMs) are the hub capability of multimodal systems: search engines rely on them for "text-to-image / image-to-text search"; content platforms use them for smart image matching, ad moderation, and image–text consistency checking; multimodal assistants use them as foundational capabilities for "chatting about images" and "asking questions about documents/screenshots." Below, we organize this layer from three angles — scenarios, principles, and models — and expand on image captioning, visual question answering, and cross-modal retrieval in subsequent subsections.From a product perspective, vision–language models (VLMs) are the hub capability of multimodal systems: search engines rely on them for "text-to-image / image-to-text search"; content platforms use them for smart image matching, ad moderation, and image–text consistency checking; multimodal assistants use them as foundational capabilities for "chatting about images" and "asking questions about documents/screenshots." Below, we organize this layer from three angles — scenarios, principles, and models — and expand on image captioning, visual question answering, and cross-modal retrieval in subsequent subsections.
The core problem is: how to map images and text into the same semantic space and perform alignment and reasoning within this space:The core problem is: how to map images and text into the same semantic space and perform alignment and reasoning within this space:
Mainstream vision–language models have roughly evolved into two categories: contrastive learning VLMs and generative multimodal large models:Mainstream vision–language models have roughly evolved into two categories: contrastive learning VLMs and generative multimodal large models:
Overall, vision–language tasks mark the point where "vision is no longer a separate perception channel" but participates together with language in higher-level knowledge representation and reasoning. Below, we expand on two directions: image captioning and visual question answering, and cross-modal retrieval and cross-modal alignment (merged into two subsections here based on content).Overall, vision–language tasks mark the point where "vision is no longer a separate perception channel" but participates together with language in higher-level knowledge representation and reasoning. Below, we expand on two directions: image captioning and visual question answering, and cross-modal retrieval and cross-modal alignment (merged into two subsections here based on content).
The goal of image captioning is to take an image as input and output a natural language description, such as "a little girl flying a kite on the grass." Traditional approaches typically used a "CNN + RNN" structure: extracting whole-image features with a convolutional network, then generating descriptions word by word with LSTM/GRU. With the emergence of Transformers and pre-trained VLMs, the mainstream paradigm has gradually shifted to an "image encoder + text decoder" structure, such as BLIP / BLIP-2, ViT + GPT, etc. For training, models are typically trained autoregressively on large-scale image–text pairs, sometimes also using reinforcement learning or contrastive losses to optimize description diversity and correctness. At the product level, image captioning is widely used for accessibility reading (generating image descriptions for screen reader software), automatically adding captions to smart albums, and providing more text indices for search systems.The goal of image captioning is to take an image as input and output a natural language description, such as "a little girl flying a kite on the grass." Traditional approaches typically used a "CNN + RNN" structure: extracting whole-image features with a convolutional network, then generating descriptions word by word with LSTM/GRU. With the emergence of Transformers and pre-trained VLMs, the mainstream paradigm has gradually shifted to an "image encoder + text decoder" structure, such as BLIP / BLIP-2, ViT + GPT, etc. For training, models are typically trained autoregressively on large-scale image–text pairs, sometimes also using reinforcement learning or contrastive losses to optimize description diversity and correctness. At the product level, image captioning is widely used for accessibility reading (generating image descriptions for screen reader software), automatically adding captions to smart albums, and providing more text indices for search systems.
Visual question answering (VQA) further brings human interaction into the picture: the model's input is no longer "image + blank prompt" but "image + question," and the output is a short answer or natural language explanation. Compared to image captioning, VQA emphasizes controllability and reasoning ability more: questions can focus on local details ("What color is the man's hat?"), relationships ("Which car is closer to the intersection?"), counting ("How many dogs are there?"), and even require external knowledge ("Which cuisine does this dish belong to?"). Early VQA models typically used an image encoder + question encoder + fusion module (e.g., bilinear pooling, attention) + classification head, outputting an answer from a limited vocabulary. Modern multimodal large models directly use an image encoder + LLM, performing natural language generation while "looking at the image," with clear advantages in open-ended answers and multi-turn dialogue.Visual question answering (VQA) further brings human interaction into the picture: the model's input is no longer "image + blank prompt" but "image + question," and the output is a short answer or natural language explanation. Compared to image captioning, VQA emphasizes controllability and reasoning ability more: questions can focus on local details ("What color is the man's hat?"), relationships ("Which car is closer to the intersection?"), counting ("How many dogs are there?"), and even require external knowledge ("Which cuisine does this dish belong to?"). Early VQA models typically used an image encoder + question encoder + fusion module (e.g., bilinear pooling, attention) + classification head, outputting an answer from a limited vocabulary. Modern multimodal large models directly use an image encoder + LLM, performing natural language generation while "looking at the image," with clear advantages in open-ended answers and multi-turn dialogue.
Both can be viewed as different "prompt templates" under a unified VLM framework:Both can be viewed as different "prompt templates" under a unified VLM framework:
+ "Describe this image in one sentence." → text;Captioning: + "Describe this image in one sentence." → text; + "Q: ... A:" → text.VQA: + "Q: ... A:" → text.Through instruction tuning, the same multimodal large model can be compatible with multiple tasks like captioning, Q&A, explanation, and tagging — this is also the foundational engineering approach for modern VLM products (multimodal assistants, image Q&A bots, etc.).Through instruction tuning, the same multimodal large model can be compatible with multiple tasks like captioning, Q&A, explanation, and tagging — this is also the foundational engineering approach for modern VLM products (multimodal assistants, image Q&A bots, etc.).
Cross-modal retrieval addresses another high-frequency need: given a piece of text, find matching images (Text-to-Image Retrieval); or given an image, find related text descriptions, product information, news reports, etc. (Image-to-Text Retrieval). These capabilities form the core of products like "text-to-image / image-to-text search," "find products by image," and "match images to news articles."Cross-modal retrieval addresses another high-frequency need: given a piece of text, find matching images (Text-to-Image Retrieval); or given an image, find related text descriptions, product information, news reports, etc. (Image-to-Text Retrieval). These capabilities form the core of products like "text-to-image / image-to-text search," "find products by image," and "match images to news articles."
The core technology is cross-modal alignment: models represented by CLIP use separate encoders for images and text (e.g., ViT and a Transformer text encoder), trained with contrastive learning on large-scale image–text pair data:The core technology is cross-modal alignment: models represented by CLIP use separate encoders for images and text (e.g., ViT and a Transformer text encoder), trained with contrastive learning on large-scale image–text pair data:
After training, simply encode all images and texts into vectors, and fast matching can be performed in the shared space through vector retrieval (nearest neighbor search):After training, simply encode all images and texts into vectors, and fast matching can be performed in the shared space through vector retrieval (nearest neighbor search):
In engineering practice, these models typically use a two-stage structure:In engineering practice, these models typically use a two-stage structure:
On the product side, cross-modal retrieval and alignment are widely used in: image search, ad retrieval (finding suitable images based on ad copy), compliance moderation (checking ad image–text consistency), content recommendation (recommending relevant images/videos based on users' reading text history), etc. With the rise of multimodal large models, these retrieval capabilities are gradually being unified into larger multimodal frameworks, provided as unified interfaces in the form of "natural language instructions + multimodal memory/vector database."On the product side, cross-modal retrieval and alignment are widely used in: image search, ad retrieval (finding suitable images based on ad copy), compliance moderation (checking ad image–text consistency), content recommendation (recommending relevant images/videos based on users' reading text history), etc. With the rise of multimodal large models, these retrieval capabilities are gradually being unified into larger multimodal frameworks, provided as unified interfaces in the form of "natural language instructions + multimodal memory/vector database."
In many businesses, the most important information is neither reflected in "objects and scenes in the image" nor in natural language descriptions of the image, but is directly written as text on the image: contract clauses, invoice amounts, street sign names, meter readings, error messages on screenshots, etc. Optical character recognition (OCR) is precisely the structured understanding task centered on "image + document layout": automatically detecting and recognizing text content from complex visual inputs, understanding document layout and structure, and further supporting search, statistics, automatic data entry, and intelligent Q&A.In many businesses, the most important information is neither reflected in "objects and scenes in the image" nor in natural language descriptions of the image, but is directly written as text on the image: contract clauses, invoice amounts, street sign names, meter readings, error messages on screenshots, etc. Optical character recognition (OCR) is precisely the structured understanding task centered on "image + document layout": automatically detecting and recognizing text content from complex visual inputs, understanding document layout and structure, and further supporting search, statistics, automatic data entry, and intelligent Q&A.
From a product perspective, OCR is the key bridge for "turning paper/image information into computable text" and is the infrastructure for digitization, automation, and intelligent office work: contract review, invoice entry, government and enterprise archive digitization, PDF-to-Word in office software, document Q&A assistants, etc., are all built on OCR capabilities. Below, we organize the OCR system from three angles — scenarios, principles, and models — and expand on core directions in subsequent subsections.From a product perspective, OCR is the key bridge for "turning paper/image information into computable text" and is the infrastructure for digitization, automation, and intelligent office work: contract review, invoice entry, government and enterprise archive digitization, PDF-to-Word in office software, document Q&A assistants, etc., are all built on OCR capabilities. Below, we organize the OCR system from three angles — scenarios, principles, and models — and expand on core directions in subsequent subsections.
The OCR system is typically divided into several key steps:The OCR system is typically divided into several key steps:
In engineering, a common combination is "specialized OCR modules + document understanding models + multimodal large models":In engineering, a common combination is "specialized OCR modules + document understanding models + multimodal large models":
Overall, OCR has evolved from early "simple character recognition" to a comprehensive document understanding system covering text + layout + structure + Q&A, and is a key pillar of enterprise digitization, government archive management, and intelligent office work. Below, we expand on three directions: text detection and recognition, document layout and table structure analysis, and document Q&A and multimodal DocVQA.Overall, OCR has evolved from early "simple character recognition" to a comprehensive document understanding system covering text + layout + structure + Q&A, and is a key pillar of enterprise digitization, government archive management, and intelligent office work. Below, we expand on three directions: text detection and recognition, document layout and table structure analysis, and document Q&A and multimodal DocVQA.
The first step of OCR is text detection: finding all text-containing regions in the input image. Street/scene text faces challenges like diverse fonts, skew and distortion, complex lighting, and severe background interference; document scenarios emphasize robust support for dense text and multi-column layouts. Methods like EAST, DBNet, etc., transform the detection problem into "pixel-level segmentation + edge learning," predicting text probability and geometric parameters on feature maps, then obtaining precise text boxes (which can be horizontal boxes or arbitrary quadrilaterals/polygons) through post-processing, balancing accuracy and speed.The first step of OCR is text detection: finding all text-containing regions in the input image. Street/scene text faces challenges like diverse fonts, skew and distortion, complex lighting, and severe background interference; document scenarios emphasize robust support for dense text and multi-column layouts. Methods like EAST, DBNet, etc., transform the detection problem into "pixel-level segmentation + edge learning," predicting text probability and geometric parameters on feature maps, then obtaining precise text boxes (which can be horizontal boxes or arbitrary quadrilaterals/polygons) through post-processing, balancing accuracy and speed.
Text recognition then crops each detected text region and converts it into a character sequence. The classic approach is represented by CRNN: first extracting features with CNN, then performing sequence modeling with RNN or Transformer, and finally outputting character sequences using CTC or attention decoding. For variable-length text, curved text, and complex languages (Chinese-English mixed, multilingual), recognition models need to simultaneously excel at visual feature modeling and character language modeling. Methods like RARE, SAR, etc., introduce spatial transformer networks (STN) or attention alignment mechanisms to correct geometric distortions and improve adaptability to complex layouts.Text recognition then crops each detected text region and converts it into a character sequence. The classic approach is represented by CRNN: first extracting features with CNN, then performing sequence modeling with RNN or Transformer, and finally outputting character sequences using CTC or attention decoding. For variable-length text, curved text, and complex languages (Chinese-English mixed, multilingual), recognition models need to simultaneously excel at visual feature modeling and character language modeling. Methods like RARE, SAR, etc., introduce spatial transformer networks (STN) or attention alignment mechanisms to correct geometric distortions and improve adaptability to complex layouts.
In engineering systems, detection and recognition typically serve as two decoupled services forming an OCR pipeline: the frontend detection breaks the image into several text lines/blocks, and the backend recognition performs character recognition on each block, optionally overlaying a language model for error correction (e.g., spell correction, number/amount validation). For specific scenarios like license plates and meter readings, specially fine-tuned detection/recognition models are used to leverage scene priors (fixed fonts, limited character sets) for higher accuracy and lower latency.In engineering systems, detection and recognition typically serve as two decoupled services forming an OCR pipeline: the frontend detection breaks the image into several text lines/blocks, and the backend recognition performs character recognition on each block, optionally overlaying a language model for error correction (e.g., spell correction, number/amount validation). For specific scenarios like license plates and meter readings, specially fine-tuned detection/recognition models are used to leverage scene priors (fixed fonts, limited character sets) for higher accuracy and lower latency.
Simply recognizing the text is not enough, especially in scenarios like long documents, reports, contracts, and invoices, where layout structure often determines the meaning and importance of information: the hierarchical relationship between titles and body text, the positioning of figures and captions, the role of headers and footers, the logical order of text inside and outside tables, etc. The goal of document layout analysis is to identify the roles and boundaries of different regions on a two-dimensional page and recover a reasonable reading order and hierarchical structure.Simply recognizing the text is not enough, especially in scenarios like long documents, reports, contracts, and invoices, where layout structure often determines the meaning and importance of information: the hierarchical relationship between titles and body text, the positioning of figures and captions, the role of headers and footers, the logical order of text inside and outside tables, etc. The goal of document layout analysis is to identify the roles and boundaries of different regions on a two-dimensional page and recover a reasonable reading order and hierarchical structure.
Models like LayoutLM / LayoutLMv2/v3, DocFormer, etc., jointly encode the content of each text token (text embedding), spatial position (bounding box coordinates), and local visual features (from CNN/ViT), modeling semantic–spatial relationships between tokens through Transformers. By training on datasets with layout annotations, the model can learn to distinguish various region types such as "title/paragraph/list/table/figure caption/header/footer" and output corresponding labels and hierarchies. These models typically serve as a "middle layer," providing structured document skeletons for contract review systems, report parsing, and archive digitization platforms.Models like LayoutLM / LayoutLMv2/v3, DocFormer, etc., jointly encode the content of each text token (text embedding), spatial position (bounding box coordinates), and local visual features (from CNN/ViT), modeling semantic–spatial relationships between tokens through Transformers. By training on datasets with layout annotations, the model can learn to distinguish various region types such as "title/paragraph/list/table/figure caption/header/footer" and output corresponding labels and hierarchies. These models typically serve as a "middle layer," providing structured document skeletons for contract review systems, report parsing, and archive digitization platforms.
Table structure recognition is a particularly critical branch of layout analysis: it not only needs to detect table regions but also further parse row/column boundaries, cell coordinates, and merged cells, ultimately reconstructing a logical table (typically represented as HTML, Markdown tables, or structured JSON with coordinates). Implementation methods include:Table structure recognition is a particularly critical branch of layout analysis: it not only needs to detect table regions but also further parse row/column boundaries, cell coordinates, and merged cells, ultimately reconstructing a logical table (typically represented as HTML, Markdown tables, or structured JSON with coordinates). Implementation methods include:
In products, these capabilities support high-value scenarios like "PDF to Word/Excel," "structured entry of receipts/invoices," and "report parsing and metric extraction," and are key components of government and enterprise office automation.In products, these capabilities support high-value scenarios like "PDF to Word/Excel," "structured entry of receipts/invoices," and "report parsing and metric extraction," and are key components of government and enterprise office automation.
When OCR and layout analysis capabilities are strong enough, the next natural demand is: no longer having people flip through documents themselves, but directly "asking the document." This is document visual question answering (DocVQA): the model answers questions on complex documents like contracts, reports, receipts, and manuals — for example, "What is the effective date of this contract?" "What is the net profit for Q4 2023 on this report page?" "Who is the buyer on this invoice?"When OCR and layout analysis capabilities are strong enough, the next natural demand is: no longer having people flip through documents themselves, but directly "asking the document." This is document visual question answering (DocVQA): the model answers questions on complex documents like contracts, reports, receipts, and manuals — for example, "What is the effective date of this contract?" "What is the net profit for Q4 2023 on this report page?" "Who is the buyer on this invoice?"
Traditional DocVQA systems are typically built as "OCR + layout model + QA head":Traditional DocVQA systems are typically built as "OCR + layout model + QA head":
With the development of multimodal large models, more and more systems are starting to directly use "document image + question" as input, letting a VLM or multimodal LLM directly generate answers or cited explanations. Under this architecture, OCR, layout, semantic understanding, and reasoning capabilities work together end-to-end within the model: the model can both see the original layout and visual cues and leverage linguistic world knowledge and reasoning patterns to complete complex question answering.With the development of multimodal large models, more and more systems are starting to directly use "document image + question" as input, letting a VLM or multimodal LLM directly generate answers or cited explanations. Under this architecture, OCR, layout, semantic understanding, and reasoning capabilities work together end-to-end within the model: the model can both see the original layout and visual cues and leverage linguistic world knowledge and reasoning patterns to complete complex question answering.
In product form, DocVQA typically appears as "contract review assistant," "invoice/report Q&A," "long document intelligent Q&A," helping users quickly locate key information from large volumes of documents, automatically generate summaries, perform clause comparison, etc., significantly reducing the burden of manual review and information retrieval.In product form, DocVQA typically appears as "contract review assistant," "invoice/report Q&A," "long document intelligent Q&A," helping users quickly locate key information from large volumes of documents, automatically generate summaries, perform clause comparison, etc., significantly reducing the burden of manual review and information retrieval.
Most of the vision capabilities introduced earlier are "discriminative": input an image, output labels, boxes, masks, or text. Another main line that has rapidly developed in recent years is generative vision: models no longer just understand images, but create or modify images, generating high-quality, multi-style visual content given text/image conditions. Image generation and editing is precisely the core capability in this direction, supporting a large number of products from AIGC drawing platforms to intelligent retouching/effects tools.Most of the vision capabilities introduced earlier are "discriminative": input an image, output labels, boxes, masks, or text. Another main line that has rapidly developed in recent years is generative vision: models no longer just understand images, but create or modify images, generating high-quality, multi-style visual content given text/image conditions. Image generation and editing is precisely the core capability in this direction, supporting a large number of products from AIGC drawing platforms to intelligent retouching/effects tools.
From a business perspective, generative vision has evolved from "technical demonstrations" to genuinely usable productivity tools: designers use it for inspiration sketches and refined drafts; marketing teams use it to batch-generate posters and ad materials; ordinary users use it to create avatars, illustrations, and wallpapers; video creators use it for cutout, background replacement, and effects. Below, we organize this layer from three angles — scenarios, principles, and models — and expand on text-to-image generation, image-to-image, and editing capabilities in subsequent subsections.From a business perspective, generative vision has evolved from "technical demonstrations" to genuinely usable productivity tools: designers use it for inspiration sketches and refined drafts; marketing teams use it to batch-generate posters and ad materials; ordinary users use it to create avatars, illustrations, and wallpapers; video creators use it for cutout, background replacement, and effects. Below, we organize this layer from three angles — scenarios, principles, and models — and expand on text-to-image generation, image-to-image, and editing capabilities in subsequent subsections.
Generative vision models primarily complete generation and editing by learning "image distributions" and "conditional control":Generative vision models primarily complete generation and editing by learning "image distributions" and "conditional control":
Current mainstream image generation and editing models are primarily diffusion models + conditional control:Current mainstream image generation and editing models are primarily diffusion models + conditional control:
At the product level, these technologies are presented to users in forms like Jimeng, Alibaba Qwen image models, FLUX, OpenAI or Gemini NanoBanana, Stable Diffusion ecosystem, Photoshop Generative Fill, Canva AI, Jianying/CapCut smart cutout and effects, gradually evolving from "toys" to formal links in the content production chain. Below, we expand on three directions: text-to-image generation, image-to-image translation, and text-driven editing.At the product level, these technologies are presented to users in forms like Jimeng, Alibaba Qwen image models, FLUX, OpenAI or Gemini NanoBanana, Stable Diffusion ecosystem, Photoshop Generative Fill, Canva AI, Jianying/CapCut smart cutout and effects, gradually evolving from "toys" to formal links in the content production chain. Below, we expand on three directions: text-to-image generation, image-to-image translation, and text-driven editing.
The core task of text-to-image generation is: given a natural language description, generate an image that matches its semantics and style as closely as possible. Modern text-to-image models are primarily based on diffusion architectures:The core task of text-to-image generation is: given a natural language description, generate an image that matches its semantics and style as closely as possible. Modern text-to-image models are primarily based on diffusion architectures:
Stable Diffusion, Imagen, DALL·E series, and other methods are trained on large-scale image–text pairs, enabling the model to both master the visual spectrum (shapes, textures, composition, lighting) and acquire a certain degree of language–vision alignment capability (understanding complex descriptions like "style," "material," "composition"). At the product level, this capability allows "people who can't draw to create images": users simply describe their ideas in natural language, and the system provides multiple visual implementations, supporting iterative exploration and refinement.Stable Diffusion, Imagen, DALL·E series, and other methods are trained on large-scale image–text pairs, enabling the model to both master the visual spectrum (shapes, textures, composition, lighting) and acquire a certain degree of language–vision alignment capability (understanding complex descriptions like "style," "material," "composition"). At the product level, this capability allows "people who can't draw to create images": users simply describe their ideas in natural language, and the system provides multiple visual implementations, supporting iterative exploration and refinement.
Text-to-image models typically support multi-style, multi-resolution output simultaneously: by incorporating style tokens, size conditions, etc., during training or inference, the same model can switch between different styles like "photorealistic, flat illustration, 3D render." Common engineering techniques include:Text-to-image models typically support multi-style, multi-resolution output simultaneously: by incorporating style tokens, size conditions, etc., during training or inference, the same model can switch between different styles like "photorealistic, flat illustration, 3D render." Common engineering techniques include:
Image-to-Image tasks, given an input image, generate another image version "constrained by it": preserving the overall structure or content of the original while achieving some transformation or enhancement. Typical forms include:Image-to-Image tasks, given an input image, generate another image version "constrained by it": preserving the overall structure or content of the original while achieving some transformation or enhancement. Typical forms include:
The key to these tasks is creating new content while preserving constraints. Diffusion models perform outstandingly here: in inpainting, the model only samples the masked region while keeping the original image unchanged in unoccluded areas, using semantic understanding and contextual information to naturally blend new content with surrounding areas in style and lighting. For style transfer, the model preserves the input structure while sampling textures and colors from the target style distribution, achieving "changing the shell without changing the bones."The key to these tasks is creating new content while preserving constraints. Diffusion models perform outstandingly here: in inpainting, the model only samples the masked region while keeping the original image unchanged in unoccluded areas, using semantic understanding and contextual information to naturally blend new content with surrounding areas in style and lighting. For style transfer, the model preserves the input structure while sampling textures and colors from the target style distribution, achieving "changing the shell without changing the bones."
In products, image-to-image capabilities support a large number of creative tools: style filters, comic conversion, one-click sky replacement, automatic beautification, old photo restoration, local retouching, etc., typically presented to users through highly visual interfaces.In products, image-to-image capabilities support a large number of creative tools: style filters, comic conversion, one-click sky replacement, automatic beautification, old photo restoration, local retouching, etc., typically presented to users through highly visual interfaces.
In traditional image editing software, users need to master a whole set of professional concepts like layers, masks, selections, and filters. Text-driven image editing attempts to replace most professional operations with natural language:In traditional image editing software, users need to master a whole set of professional concepts like layers, masks, selections, and filters. Text-driven image editing attempts to replace most professional operations with natural language:
Technically, text-driven editing is typically built on top of text-to-image diffusion models, implemented through several approaches:Technically, text-driven editing is typically built on top of text-to-image diffusion models, implemented through several approaches:
Jimeng, FLUX, Alibaba Qwen image models, the Stable Diffusion ecosystem, Canva AI, and other products all provide similar capabilities: users can complete complex edits through simple text and minimal interaction. For professional users, this becomes a "smart assistant" that accelerates the creative workflow; for ordinary users, it greatly lowers the barrier to image editing.Jimeng, FLUX, Alibaba Qwen image models, the Stable Diffusion ecosystem, Canva AI, and other products all provide similar capabilities: users can complete complex edits through simple text and minimal interaction. For professional users, this becomes a "smart assistant" that accelerates the creative workflow; for ordinary users, it greatly lowers the barrier to image editing.
In tasks like low-level vision enhancement, compression encoding, and image generation and editing, we often need to answer a seemingly subjective question: "Does this image look good?" Manual inspection clearly cannot scale, and traditional metrics like PSNR often do not align with human subjective perception. The goal of image quality assessment (IQA) is to establish an automated mechanism for scoring or ranking the subjective/objective quality of images, serving as the key link between "low-level algorithm output" and "real user experience."In tasks like low-level vision enhancement, compression encoding, and image generation and editing, we often need to answer a seemingly subjective question: "Does this image look good?" Manual inspection clearly cannot scale, and traditional metrics like PSNR often do not align with human subjective perception. The goal of image quality assessment (IQA) is to establish an automated mechanism for scoring or ranking the subjective/objective quality of images, serving as the key link between "low-level algorithm output" and "real user experience."
From a system perspective, IQA is the "gatekeeper" and "tuning reference" in many pipelines: e-commerce/content platforms use it to filter out blurry, noisy, heavily compressed uploaded images; phone cameras/albums use it to pick the "best shot" from a burst; cloud-based enhancement and compression services use it for before-and-after comparison evaluation to guide model iteration. Below, we organize IQA from three dimensions — scenarios, principles, and models — and expand on evaluation types and metrics/learning paradigms in subsequent subsections.From a system perspective, IQA is the "gatekeeper" and "tuning reference" in many pipelines: e-commerce/content platforms use it to filter out blurry, noisy, heavily compressed uploaded images; phone cameras/albums use it to pick the "best shot" from a burst; cloud-based enhancement and compression services use it for before-and-after comparison evaluation to guide model iteration. Below, we organize IQA from three dimensions — scenarios, principles, and models — and expand on evaluation types and metrics/learning paradigms in subsequent subsections.
The core of IQA is to characterize image quality from two dimensions: the degree of distortion relative to a reference image and the goodness of human subjective perception:The core of IQA is to characterize image quality from two dimensions: the degree of distortion relative to a reference image and the goodness of human subjective perception:
IQA models are roughly divided into two categories: traditional hand-crafted feature metrics and deep learning-based quality prediction:IQA models are roughly divided into two categories: traditional hand-crafted feature metrics and deep learning-based quality prediction:
Overall, IQA is not a single metric where "higher is always better," but an evaluation system related to specific business objectives: in some scenarios (e.g., surveillance enhancement), preserving detail and recognizability is more important than visual naturalness; in content creation platforms, subjective perception and aesthetic standards dominate. Therefore, common industry practice is: building on top of general IQA models, fine-tuning or learning weights with a small amount of business data to construct "task-aware" quality evaluators.Overall, IQA is not a single metric where "higher is always better," but an evaluation system related to specific business objectives: in some scenarios (e.g., surveillance enhancement), preserving detail and recognizability is more important than visual naturalness; in content creation platforms, subjective perception and aesthetic standards dominate. Therefore, common industry practice is: building on top of general IQA models, fine-tuning or learning weights with a small amount of business data to construct "task-aware" quality evaluators.
Based on whether a high-quality reference image exists, IQA can be divided into three categories: full-reference (FR-IQA), no-reference (NR-IQA), and pseudo-reference.Based on whether a high-quality reference image exists, IQA can be divided into three categories: full-reference (FR-IQA), no-reference (NR-IQA), and pseudo-reference.
In full-reference IQA, we assume the existence of an ideal high-quality reference image, and the image under evaluation is its degraded version after compression, transmission, or processing. The model quantifies the degree of distortion by comparing the two pixel-by-pixel or at the feature level. PSNR is the simplest metric (based on mean squared error); SSIM/MS-SSIM/FSIM, etc., further consider brightness, contrast, structure, or phase information, coming closer to human visual perception to some extent. These metrics are very suitable for evaluating methods like encoding/decoding, super-resolution, and denoising during the algorithm development phase, but in real business, reference images are often unavailable, limiting application scenarios.In full-reference IQA, we assume the existence of an ideal high-quality reference image, and the image under evaluation is its degraded version after compression, transmission, or processing. The model quantifies the degree of distortion by comparing the two pixel-by-pixel or at the feature level. PSNR is the simplest metric (based on mean squared error); SSIM/MS-SSIM/FSIM, etc., further consider brightness, contrast, structure, or phase information, coming closer to human visual perception to some extent. These metrics are very suitable for evaluating methods like encoding/decoding, super-resolution, and denoising during the algorithm development phase, but in real business, reference images are often unavailable, limiting application scenarios.
No-reference IQA (Blind IQA) is the more common setting in real systems: only the image under evaluation itself is available, with no reference at all. Early no-reference methods (e.g., BRISQUE, NIQE, BLIINDS, etc.) were mainly based on natural scene statistics: assuming that high-quality natural images have stable forms in certain statistical distributions, and that distortion causes changes in statistical features, allowing models to be trained to predict quality scores based on these features. In the deep learning era, NR-IQA models typically directly use CNN / ViT to extract features and regress quality scores or learn ranking relationships on datasets with human subjective ratings (MOS), enabling them to cover various distortion types like noise, blur, compression artifacts, and exposure anomalies.No-reference IQA (Blind IQA) is the more common setting in real systems: only the image under evaluation itself is available, with no reference at all. Early no-reference methods (e.g., BRISQUE, NIQE, BLIINDS, etc.) were mainly based on natural scene statistics: assuming that high-quality natural images have stable forms in certain statistical distributions, and that distortion causes changes in statistical features, allowing models to be trained to predict quality scores based on these features. In the deep learning era, NR-IQA models typically directly use CNN / ViT to extract features and regress quality scores or learn ranking relationships on datasets with human subjective ratings (MOS), enabling them to cover various distortion types like noise, blur, compression artifacts, and exposure anomalies.
Pseudo-reference / reduced-reference IQA falls between the two: without a truly high-quality reference, using some obtainable approximate version (e.g., pre-compression low-resolution image, model-predicted "clean image") as a reference to estimate the degree of degradation. This approach is common in online video quality monitoring and encoding/decoding optimization tasks, striking a balance between cost and accuracy.Pseudo-reference / reduced-reference IQA falls between the two: without a truly high-quality reference, using some obtainable approximate version (e.g., pre-compression low-resolution image, model-predicted "clean image") as a reference to estimate the degree of degradation. This approach is common in online video quality monitoring and encoding/decoding optimization tasks, striking a balance between cost and accuracy.
At the specific implementation level, IQA employs various metrics and learning paradigms to approximate human subjective perception.At the specific implementation level, IQA employs various metrics and learning paradigms to approximate human subjective perception.
Traditional metrics:Traditional metrics:
Perceptual metrics: LPIPS, DISTS, etc., compute vector differences in the internal feature layers of pre-trained deep networks (VGG, AlexNet, ViT, etc.), weighted by the importance of different layers, yielding a "distance in feature space" that has higher correlation with subjective perceptual similarity. They are particularly suitable as training objectives or evaluation metrics for generative tasks (super-resolution, generation, editing), used to measure "how similar it looks."Perceptual metrics: LPIPS, DISTS, etc., compute vector differences in the internal feature layers of pre-trained deep networks (VGG, AlexNet, ViT, etc.), weighted by the importance of different layers, yielding a "distance in feature space" that has higher correlation with subjective perceptual similarity. They are particularly suitable as training objectives or evaluation metrics for generative tasks (super-resolution, generation, editing), used to measure "how similar it looks."
Learning-based quality prediction: deep NR-IQA models (e.g., RankIQA, DBCNN, HyperIQA, MUSIQ, etc.) directly score or rank images:Learning-based quality prediction: deep NR-IQA models (e.g., RankIQA, DBCNN, HyperIQA, MUSIQ, etc.) directly score or rank images:
With the widespread adoption of large-scale pre-trained vision models, more and more IQA methods adopt the "pre-trained backbone + lightweight head" paradigm: leveraging rich visual representations from CLIP, ViT, etc., and fine-tuning on relatively little MOS data, thereby maintaining good generalization across distortion types and scenarios.With the widespread adoption of large-scale pre-trained vision models, more and more IQA methods adopt the "pre-trained backbone + lightweight head" paradigm: leveraging rich visual representations from CLIP, ViT, etc., and fine-tuning on relatively little MOS data, thereby maintaining good generalization across distortion types and scenarios.
In engineering deployment, multiple of the above metrics are typically combined: for example, FR-IQA metrics for evaluating algorithm improvements during the experimental phase; deep NR-IQA models for online real-time quality inspection; perceptual metrics for internal optimization of generative tasks. Through A/B experiments, these automatic metrics are aligned with real user data (click-through rate, completion rate, complaint rate, etc.), gradually building a "perceptual quality measurement system" highly correlated with business objectives.In engineering deployment, multiple of the above metrics are typically combined: for example, FR-IQA metrics for evaluating algorithm improvements during the experimental phase; deep NR-IQA models for online real-time quality inspection; perceptual metrics for internal optimization of generative tasks. Through A/B experiments, these automatic metrics are aligned with real user data (click-through rate, completion rate, complaint rate, etc.), gradually building a "perceptual quality measurement system" highly correlated with business objectives.
As applications move from "flat images/video" to autonomous driving, robotics, AR/VR/XR, and similar scenarios, systems are no longer satisfied with merely looking at "2D pixels" — they need to understand the three-dimensional structure, scale, and pose relationships of the real world. These tasks are collectively referred to as the 3D / spatial modality: they encompass both precise geometric and topological modeling, as well as semantic understanding, localization and navigation, and content generation within 3D space. On one end, they connect to various sensors such as LiDAR, RGB‑D, and IMU; on the other, they connect to autonomous driving perception modules, robot navigation systems, ARKit/ARCore environment models, mobile 3D scanning and modeling applications, and digital twin platforms.As applications move from "flat images/video" to autonomous driving, robotics, AR/VR/XR, and similar scenarios, systems are no longer satisfied with merely looking at "2D pixels" — they need to understand the three-dimensional structure, scale, and pose relationships of the real world. These tasks are collectively referred to as the 3D / spatial modality: they encompass both precise geometric and topological modeling, as well as semantic understanding, localization and navigation, and content generation within 3D space. On one end, they connect to various sensors such as LiDAR, RGB‑D, and IMU; on the other, they connect to autonomous driving perception modules, robot navigation systems, ARKit/ARCore environment models, mobile 3D scanning and modeling applications, and digital twin platforms.
In 2D vision, we only see "the world after it has been photographed"; but in scenarios like autonomous driving, robotics, and AR/VR, what matters more is: the position, shape, and structure of the real world in 3D space. 3D perception and reconstruction aims to recover the three-dimensional geometric information of the environment from multiple sensors (cameras, LiDAR, depth cameras, etc.) and express it in forms such as point clouds, voxels, meshes, and implicit fields, providing the foundation for path planning, physics simulation, digital twins, and 3D content generation.In 2D vision, we only see "the world after it has been photographed"; but in scenarios like autonomous driving, robotics, and AR/VR, what matters more is: the position, shape, and structure of the real world in 3D space. 3D perception and reconstruction aims to recover the three-dimensional geometric information of the environment from multiple sensors (cameras, LiDAR, depth cameras, etc.) and express it in forms such as point clouds, voxels, meshes, and implicit fields, providing the foundation for path planning, physics simulation, digital twins, and 3D content generation.
In engineering practice, this layer covers multiple technical directions from point cloud processing to multi-view geometric reconstruction to neural radiance fields / neural field rendering, corresponding to product forms such as autonomous driving 3D perception modules, ARKit/ARCore environment modeling, mobile 3D scanning/modeling apps, and digital twin city/campus modeling platforms. The following expands from three angles — scenarios, principles, and models — and further subdivides into several key sub-directions.In engineering practice, this layer covers multiple technical directions from point cloud processing to multi-view geometric reconstruction to neural radiance fields / neural field rendering, corresponding to product forms such as autonomous driving 3D perception modules, ARKit/ARCore environment modeling, mobile 3D scanning/modeling apps, and digital twin city/campus modeling platforms. The following expands from three angles — scenarios, principles, and models — and further subdivides into several key sub-directions.
Starting from this layer, traditional geometry, deep learning, implicit representations, and explicit meshes are closely intertwined — the goal is both to solve "how to accurately reconstruct the real world" and to balance real-time performance and usability, serving higher-level 3D scene understanding, 3D generation, and editing.Starting from this layer, traditional geometry, deep learning, implicit representations, and explicit meshes are closely intertwined — the goal is both to solve "how to accurately reconstruct the real world" and to balance real-time performance and usability, serving higher-level 3D scene understanding, 3D generation, and editing.
For autonomous driving, robotics, and high-precision surveying, LiDAR point clouds are one of the most critical forms of 3D sensing information. A point cloud is a sparse set of points composed of 3D coordinates (sometimes accompanied by reflection intensity, timestamps, etc.), lacking a regular grid structure, which poses challenges for traditional convolution. The goal of point cloud processing is to extract useful geometric and semantic information from these unstructured points — for example, "this is a car," "this is a curb/ground," "this is a building."For autonomous driving, robotics, and high-precision surveying, LiDAR point clouds are one of the most critical forms of 3D sensing information. A point cloud is a sparse set of points composed of 3D coordinates (sometimes accompanied by reflection intensity, timestamps, etc.), lacking a regular grid structure, which poses challenges for traditional convolution. The goal of point cloud processing is to extract useful geometric and semantic information from these unstructured points — for example, "this is a car," "this is a curb/ground," "this is a building."
In point cloud classification and segmentation tasks, we typically focus on: which category a given point (or point cluster) belongs to, such as car, pedestrian, ground, curb, building, vegetation, etc., or performing semantic/instance segmentation on the scene. From a modeling perspective, approaches can be roughly divided into three categories:In point cloud classification and segmentation tasks, we typically focus on: which category a given point (or point cluster) belongs to, such as car, pedestrian, ground, curb, building, vegetation, etc., or performing semantic/instance segmentation on the scene. From a modeling perspective, approaches can be roughly divided into three categories:
In 3D object detection, the goal is no longer simply labeling points, but predicting 3D bounding boxes (position, size, orientation) and their categories — this is the core of autonomous driving environment perception. Typical methods such as VoxelNet, SECOND, PointPillars, and CenterPoint usually convert point clouds into voxel or pillar representations and perform detection regression in BEV or 3D space. Methods like CenterPoint use a "center point detection" paradigm, directly detecting object centers and their sizes/orientations on BEV, balancing accuracy and speed. As deep learning and sensor hardware evolve, 3D detection can now achieve real-time inference on automotive-grade chips, becoming one of the foundational modules of the autonomous driving perception stack.In 3D object detection, the goal is no longer simply labeling points, but predicting 3D bounding boxes (position, size, orientation) and their categories — this is the core of autonomous driving environment perception. Typical methods such as VoxelNet, SECOND, PointPillars, and CenterPoint usually convert point clouds into voxel or pillar representations and perform detection regression in BEV or 3D space. Methods like CenterPoint use a "center point detection" paradigm, directly detecting object centers and their sizes/orientations on BEV, balancing accuracy and speed. As deep learning and sensor hardware evolve, 3D detection can now achieve real-time inference on automotive-grade chips, becoming one of the foundational modules of the autonomous driving perception stack.
Without LiDAR, can we still "understand" 3D? The answer is yes — multi-view geometry and 3D reconstruction rely on "multiple photos + camera motion." By capturing the same scene from different viewpoints, we can use geometric constraints to recover camera poses and spatial structure — this is the classic SfM/MVS pipeline.Without LiDAR, can we still "understand" 3D? The answer is yes — multi-view geometry and 3D reconstruction rely on "multiple photos + camera motion." By capturing the same scene from different viewpoints, we can use geometric constraints to recover camera poses and spatial structure — this is the classic SfM/MVS pipeline.
SfM (Structure‑from‑Motion) primarily solves two problems:SfM (Structure‑from‑Motion) primarily solves two problems:
Typical tools such as COLMAP and OpenMVG, through feature extraction and matching (SIFT/ORB, etc.) and incremental or global BA (Bundle Adjustment), can automatically recover sparse point clouds and camera poses from uncalibrated image collections.Typical tools such as COLMAP and OpenMVG, through feature extraction and matching (SIFT/ORB, etc.) and incremental or global BA (Bundle Adjustment), can automatically recover sparse point clouds and camera poses from uncalibrated image collections.
Building on this, MVS (Multi‑View Stereo) uses multi-view photometric consistency to generate dense point clouds: performing depth estimation for each pixel/ray, gradually filling in the geometric details of the scene.Building on this, MVS (Multi‑View Stereo) uses multi-view photometric consistency to generate dense point clouds: performing depth estimation for each pixel/ray, gradually filling in the geometric details of the scene.
After obtaining a dense point cloud, the next step is mesh reconstruction:After obtaining a dense point cloud, the next step is mesh reconstruction:
In terms of product form, this entire pipeline has been delivered through desktop software, cloud services, and SDKs. For example: 3D scanning apps on phones call SfM/MVS-like processes in the background, automatically outputting a mesh model that can be imported into game engines after the user "circles around and takes photos" or "scans a video"; digital twin platforms run large-scale reconstruction at the city/campus scale using aerial imagery + street view data to generate interactive 3D scenes.In terms of product form, this entire pipeline has been delivered through desktop software, cloud services, and SDKs. For example: 3D scanning apps on phones call SfM/MVS-like processes in the background, automatically outputting a mesh model that can be imported into game engines after the user "circles around and takes photos" or "scans a video"; digital twin platforms run large-scale reconstruction at the city/campus scale using aerial imagery + street view data to generate interactive 3D scenes.
Traditional SfM/MVS/mesh reconstruction can produce well-structured explicit geometry, but still has limitations in rendering quality, viewpoint continuity, and detail representation; neural radiance fields (NeRF) and its follow-up work have redefined 3D reconstruction and novel view synthesis through implicit fields + volume rendering.Traditional SfM/MVS/mesh reconstruction can produce well-structured explicit geometry, but still has limitations in rendering quality, viewpoint continuity, and detail representation; neural radiance fields (NeRF) and its follow-up work have redefined 3D reconstruction and novel view synthesis through implicit fields + volume rendering.
In NeRF, the entire 3D scene is modeled as a continuous function:In NeRF, the entire 3D scene is modeled as a continuous function:
$$$$
F_\theta(\mathbf{x}, \mathbf{d}) = (\sigma, \mathbf{c})F_\theta(\mathbf{x}, \mathbf{d}) = (\sigma, \mathbf{c})
$$$$
where $\mathbf{x}$ represents the position of a point in 3D space, $\mathbf{d}$ represents the viewing direction, $\sigma$ represents volume density, $\mathbf{c}$ represents color, and $\theta$ represents network parameters.where $\mathbf{x}$ represents the position of a point in 3D space, $\mathbf{d}$ represents the viewing direction, $\sigma$ represents volume density, $\mathbf{c}$ represents color, and $\theta$ represents network parameters.
Given a point position x and viewing direction d in 3D space, the network outputs the corresponding volume density σ and color c at that point. By performing volume rendering integration along the camera ray direction on this mapping function, we can obtain the pixel color for that camera pose; conversely, given a set of multi-view photos and their camera parameters, we can solve for the model parameters θ by minimizing the error between rendered results and real images. Once the model is trained, simply changing the camera pose allows synthesizing novel view images that were "never actually photographed" (Novel View Synthesis).Given a point position x and viewing direction d in 3D space, the network outputs the corresponding volume density σ and color c at that point. By performing volume rendering integration along the camera ray direction on this mapping function, we can obtain the pixel color for that camera pose; conversely, given a set of multi-view photos and their camera parameters, we can solve for the model parameters θ by minimizing the error between rendered results and real images. Once the model is trained, simply changing the camera pose allows synthesizing novel view images that were "never actually photographed" (Novel View Synthesis).
Traditional NeRF has relatively slow training and rendering speeds; subsequent work such as Instant‑NGP uses techniques like multi-resolution hash grid encoding to dramatically accelerate convergence and inference; Gaussian Splatting replaces scene representation with 3D Gaussian particles, achieving high-quality, real-time novel view rendering through efficient rasterization strategies. Meanwhile, a large body of work has extended NeRF/Gaussian with editability, multimodality, and composability, gradually moving from research prototypes to engineering systems.Traditional NeRF has relatively slow training and rendering speeds; subsequent work such as Instant‑NGP uses techniques like multi-resolution hash grid encoding to dramatically accelerate convergence and inference; Gaussian Splatting replaces scene representation with 3D Gaussian particles, achieving high-quality, real-time novel view rendering through efficient rasterization strategies. Meanwhile, a large body of work has extended NeRF/Gaussian with editability, multimodality, and composability, gradually moving from research prototypes to engineering systems.
At the productization level, NeRF/Gaussian technologies have been embedded into various 3D AI products:At the productization level, NeRF/Gaussian technologies have been embedded into various 3D AI products:
If 3D perception and reconstruction answers "what does this world look like," then 3D scene understanding and localization further answers: "Where am I in this world? Which places in this world are traversable, and which are obstacles?" For robot vacuums, AGV robots, drones, AR navigation, and indoor positioning systems, the ability to self-localize, self-map, and autonomously plan paths in a 3D environment is a prerequisite for survival.If 3D perception and reconstruction answers "what does this world look like," then 3D scene understanding and localization further answers: "Where am I in this world? Which places in this world are traversable, and which are obstacles?" For robot vacuums, AGV robots, drones, AR navigation, and indoor positioning systems, the ability to self-localize, self-map, and autonomously plan paths in a 3D environment is a prerequisite for survival.
This body of work primarily revolves around 3D semantic understanding and SLAM (Simultaneous Localization and Mapping): the former performs semantic segmentation and traversable region identification within reconstructed 3D scenes, while the latter uses visual/IMU/LiDAR and other sensors for camera/robot pose estimation and map construction. In engineering, this layer is typically embedded as SDKs or algorithm modules into robot chassis, drone flight controllers, or mobile AR engines.This body of work primarily revolves around 3D semantic understanding and SLAM (Simultaneous Localization and Mapping): the former performs semantic segmentation and traversable region identification within reconstructed 3D scenes, while the latter uses visual/IMU/LiDAR and other sensors for camera/robot pose estimation and map construction. In engineering, this layer is typically embedded as SDKs or algorithm modules into robot chassis, drone flight controllers, or mobile AR engines.
Overall, 3D scene understanding and localization form the foundation for robots to "get moving": they must both build a reliable self-localization framework in complex 3D worlds and make maps "meaningful," thereby supporting high-level task planning and human-robot interaction.Overall, 3D scene understanding and localization form the foundation for robots to "get moving": they must both build a reliable self-localization framework in complex 3D worlds and make maps "meaningful," thereby supporting high-level task planning and human-robot interaction.
In a purely geometric map, all structures are just undifferentiated points/voxels; in real applications, what we care about is: where is the ground, where are the walls, where are tables or shelves, and where is traversable. 3D semantic segmentation aims to assign a semantic label to every point or voxel, transforming "pure geometry" into "geometry + semantics."In a purely geometric map, all structures are just undifferentiated points/voxels; in real applications, what we care about is: where is the ground, where are the walls, where are tables or shelves, and where is traversable. 3D semantic segmentation aims to assign a semantic label to every point or voxel, transforming "pure geometry" into "geometry + semantics."
In indoor/outdoor scenes, typical targets include:In indoor/outdoor scenes, typical targets include:
In terms of modeling, 3D semantic segmentation commonly uses:In terms of modeling, 3D semantic segmentation commonly uses:
In applications such as robot vacuums and AGV robots, semantic segmentation results are further abstracted into semantic maps: for example, dividing rooms into bedroom/living room/kitchen, and dividing warehouse spaces into shelf areas/aisles/restricted zones. Robots not only know "where they can go" but can also tailor different strategies based on room type (e.g., avoiding carpeted areas in bedrooms, prioritizing certain shelf zones in warehouses).In applications such as robot vacuums and AGV robots, semantic segmentation results are further abstracted into semantic maps: for example, dividing rooms into bedroom/living room/kitchen, and dividing warehouse spaces into shelf areas/aisles/restricted zones. Robots not only know "where they can go" but can also tailor different strategies based on room type (e.g., avoiding carpeted areas in bedrooms, prioritizing certain shelf zones in warehouses).
The goal of SLAM (Simultaneous Localization and Mapping) is: in an unknown environment, to estimate one's own trajectory while moving and simultaneously build a map of the environment. For indoor environments without high-precision external positioning (such as RTK‑GNSS), SLAM is the preferred solution for the vast majority of robots and AR engines.The goal of SLAM (Simultaneous Localization and Mapping) is: in an unknown environment, to estimate one's own trajectory while moving and simultaneously build a map of the environment. For indoor environments without high-precision external positioning (such as RTK‑GNSS), SLAM is the preferred solution for the vast majority of robots and AR engines.
In visual SLAM, methods represented by ORB‑SLAM, DSO, and VINS‑Mono/VINS‑Fusion are typically divided into several key modules:In visual SLAM, methods represented by ORB‑SLAM, DSO, and VINS‑Mono/VINS‑Fusion are typically divided into several key modules:
Pure vision tends to fail in texture-poor areas or under drastic lighting changes, so in practice multi-sensor fusion localization is generally adopted:Pure vision tends to fail in texture-poor areas or under drastic lighting changes, so in practice multi-sensor fusion localization is generally adopted:
At the product level, these methods are typically encapsulated as part of robot chassis controllers, drone flight controllers, AR engines (such as Visual‑Inertial SLAM in ARKit/ARCore), or indoor positioning SDKs, shielding upper-layer applications from complex state estimation and graph optimization logic, allowing developers to directly obtain "real-time pose + map."At the product level, these methods are typically encapsulated as part of robot chassis controllers, drone flight controllers, AR engines (such as Visual‑Inertial SLAM in ARKit/ARCore), or indoor positioning SDKs, shielding upper-layer applications from complex state estimation and graph optimization logic, allowing developers to directly obtain "real-time pose + map."
With stable pose estimation and geometric/semantic maps in place, the next step is to make the robot "move intelligently." This part primarily involves semantic map construction, path planning, and obstacle avoidance.With stable pose estimation and geometric/semantic maps in place, the next step is to make the robot "move intelligently." This part primarily involves semantic map construction, path planning, and obstacle avoidance.
AR navigation and indoor positioning systems also essentially rely on similar semantic maps and path planning, except the "executor" changes from a robot to a person: the system obtains the user's device pose through SLAM, plans a walking path on the semantic map, and then visualizes the path as an augmented reality overlay on the real-world view.AR navigation and indoor positioning systems also essentially rely on similar semantic maps and path planning, except the "executor" changes from a robot to a person: the system obtains the user's device pose through SLAM, plans a walking path on the semantic map, and then visualizes the path as an augmented reality overlay on the real-world view.
If 3D perception and SLAM are about "capturing and understanding" geometry from the real world, then 3D generation and editing take the perspective of content production: how to use AI to automatically produce and modify 3D assets. This directly addresses the enormous content demands of gaming, film, digital humans, virtual spaces, e-commerce displays, 3D printing, and more.If 3D perception and SLAM are about "capturing and understanding" geometry from the real world, then 3D generation and editing take the perspective of content production: how to use AI to automatically produce and modify 3D assets. This directly addresses the enormous content demands of gaming, film, digital humans, virtual spaces, e-commerce displays, 3D printing, and more.
In the past two to three years, with breakthroughs in technologies such as NeRF/Gaussian, SDF representations, and multimodal diffusion models, 3D generation has entered a period of rapid development: one-click generation of 3D models or scenes from text, images, and video has become a reality, and major cloud providers and startup teams have launched products such as "Hunyuan 3D," Tripo, and the DreamFusion / Magic3D series of methods, implemented as online tools, gradually moving 3D production toward "accessible to everyone." 3D generation and editing can be roughly divided into four capabilities: text-to-3D, image/video-to-3D, model optimization and editing, and rigging and animation.In the past two to three years, with breakthroughs in technologies such as NeRF/Gaussian, SDF representations, and multimodal diffusion models, 3D generation has entered a period of rapid development: one-click generation of 3D models or scenes from text, images, and video has become a reality, and major cloud providers and startup teams have launched products such as "Hunyuan 3D," Tripo, and the DreamFusion / Magic3D series of methods, implemented as online tools, gradually moving 3D production toward "accessible to everyone." 3D generation and editing can be roughly divided into four capabilities: text-to-3D, image/video-to-3D, model optimization and editing, and rigging and animation.
In this layer, traditional 3D DCC (Maya/Blender/3ds Max, etc.) and AI toolchains are gradually merging: many 3D AI services are embedded into existing production workflows as plugins or cloud interfaces, allowing modelers/artists to rapidly iterate assets through human-AI collaboration.In this layer, traditional 3D DCC (Maya/Blender/3ds Max, etc.) and AI toolchains are gradually merging: many 3D AI services are embedded into existing production workflows as plugins or cloud interfaces, allowing modelers/artists to rapidly iterate assets through human-AI collaboration.
The goal of Text‑to‑3D is: given a natural language description, such as "a cartoon-style yellow duck toy with a blue scarf, suitable for children's toy display," the system automatically generates an editable 3D model (Mesh/NeRF/SDF/Gaussian, etc.). This is a classic application of combining large language models / multimodal models with 3D representations.The goal of Text‑to‑3D is: given a natural language description, such as "a cartoon-style yellow duck toy with a blue scarf, suitable for children's toy display," the system automatically generates an editable 3D model (Mesh/NeRF/SDF/Gaussian, etc.). This is a classic application of combining large language models / multimodal models with 3D representations.
Typical technical paths include:Typical technical paths include:
At the scene level, scene rough model capabilities allow users to describe spatial layouts using natural language or rough sketches — for example, "a living room with floor-to-ceiling windows, an L-shaped sofa on the left, a coffee table in the middle, and bookshelves and a TV stand on the right" — and the system automatically builds a geometrically and semantically reasonable 3D layout sketch. Subsequently, models and materials can be refined in DCC tools, or usable scene prototypes can be quickly produced directly through the "scene generation" capabilities in tools like Hunyuan 3D and Tripo.At the scene level, scene rough model capabilities allow users to describe spatial layouts using natural language or rough sketches — for example, "a living room with floor-to-ceiling windows, an L-shaped sofa on the left, a coffee table in the middle, and bookshelves and a TV stand on the right" — and the system automatically builds a geometrically and semantically reasonable 3D layout sketch. Subsequently, models and materials can be refined in DCC tools, or usable scene prototypes can be quickly produced directly through the "scene generation" capabilities in tools like Hunyuan 3D and Tripo.
Currently, multiple platforms have launched Text‑to‑3D products for designers and developers:Currently, multiple platforms have launched Text‑to‑3D products for designers and developers:
Compared to pure text, generating 3D models from images or video provides stronger geometric constraints and better visual consistency. Therefore, a large number of 3D AI products support image-to-3D / video-to-3D:Compared to pure text, generating 3D models from images or video provides stronger geometric constraints and better visual consistency. Therefore, a large number of 3D AI products support image-to-3D / video-to-3D:
Generating 3D geometry is only the first step; substantial model optimization and editing work is still needed afterward:Generating 3D geometry is only the first step; substantial model optimization and editing work is still needed afterward:
Products such as Hunyuan 3D and Tripo often connect the above workflows end-to-end: users start from photos/videos or simple text, and the system internally completes reconstruction, retopology, texturing, and export, allowing non-professional users to obtain "plug-and-play" 3D models within minutes, dramatically shortening the time from concept to asset.Products such as Hunyuan 3D and Tripo often connect the above workflows end-to-end: users start from photos/videos or simple text, and the system internally completes reconstruction, retopology, texturing, and export, allowing non-professional users to obtain "plug-and-play" 3D models within minutes, dramatically shortening the time from concept to asset.
Static models are only half the content; 3D assets that "can move" are more critical in gaming, film, virtual humans, and interactive applications. This involves skeletal rigging, weight painting, animation, and physics simulation — traditionally high-barrier professional tasks that are now increasingly assisted or even semi-automated by AI tools.Static models are only half the content; 3D assets that "can move" are more critical in gaming, film, virtual humans, and interactive applications. This involves skeletal rigging, weight painting, animation, and physics simulation — traditionally high-barrier professional tasks that are now increasingly assisted or even semi-automated by AI tools.
In terms of products and ecosystems, these capabilities are often embedded in:In terms of products and ecosystems, these capabilities are often embedded in:
As 3D generation and editing technologies mature, the entire 3D content production pipeline is evolving from "centered on professional DCC tools" to "AI-driven human-AI collaboration": AI handles generation and extensive foundational work, while humans make more decisions on style definition, quality control, and key design nodes. Hunyuan 3D, Tripo, and other next-generation 3D AI products are a concentrated manifestation of this trend, providing faster and more accessible 3D infrastructure for upstream gaming, film, AR/VR, digital twin, and virtual human applications.As 3D generation and editing technologies mature, the entire 3D content production pipeline is evolving from "centered on professional DCC tools" to "AI-driven human-AI collaboration": AI handles generation and extensive foundational work, while humans make more decisions on style definition, quality control, and key design nodes. Hunyuan 3D, Tripo, and other next-generation 3D AI products are a concentrated manifestation of this trend, providing faster and more accessible 3D infrastructure for upstream gaming, film, AR/VR, digital twin, and virtual human applications.
In the overall technology stack, "audio" corresponds to the perception and generation of acoustic signals: it includes not only the processing of raw waveforms and spectra, but also converting speech to text, understanding "who is speaking" and "what is being said," as well as further creating and synthesizing sound and music. Similar to vision, audio can also be broken down into multiple layers: the bottom layer of waveform and spectrum processing is responsible for "hearing clearly"; the middle layer of speech recognition and speaker technology is responsible for "understanding who is saying what"; above that are more abstract layers of audio/music understanding and speech and music generation. This entire set of capabilities collectively supports products such as real-time meeting captions, voice assistants, podcast post-production tuning, smart speakers, acoustic security monitoring, and music recommendation and generation.In the overall technology stack, "audio" corresponds to the perception and generation of acoustic signals: it includes not only the processing of raw waveforms and spectra, but also converting speech to text, understanding "who is speaking" and "what is being said," as well as further creating and synthesizing sound and music. Similar to vision, audio can also be broken down into multiple layers: the bottom layer of waveform and spectrum processing is responsible for "hearing clearly"; the middle layer of speech recognition and speaker technology is responsible for "understanding who is saying what"; above that are more abstract layers of audio/music understanding and speech and music generation. This entire set of capabilities collectively supports products such as real-time meeting captions, voice assistants, podcast post-production tuning, smart speakers, acoustic security monitoring, and music recommendation and generation.
At the lowest level of audio technology, what we first care about is not "what is being said," "who is speaking," or "what style of music this is," but rather whether the sound itself is clean and clear enough to hear. This layer primarily works at the waveform and spectrum level, using operations such as resampling, enhancement, noise reduction, and separation to process noisy, distorted, and mixed raw audio into "clean signals" that are more suitable for subsequent recognition, analysis, and generation. It can be likened to "image enhancement + denoising + foreground/background separation" in vision — it is more about performing acoustic-level cleanup rather than directly processing semantics.At the lowest level of audio technology, what we first care about is not "what is being said," "who is speaking," or "what style of music this is," but rather whether the sound itself is clean and clear enough to hear. This layer primarily works at the waveform and spectrum level, using operations such as resampling, enhancement, noise reduction, and separation to process noisy, distorted, and mixed raw audio into "clean signals" that are more suitable for subsequent recognition, analysis, and generation. It can be likened to "image enhancement + denoising + foreground/background separation" in vision — it is more about performing acoustic-level cleanup rather than directly processing semantics.
From a product perspective, this layer is almost "invisible" behind every audio product: real-time noise reduction in meeting software, post-production tuning for podcasts and short videos, the "voice enhancement mode" in voice recorders and phones, the "beautify voice" toggle on live streaming platforms, and the front-end preprocessing for ASR/speaker verification models — all of these are direct manifestations of waveform-level audio processing. Below, we continue to organize this from three angles — scenarios, principles, and models — and in subsequent subsections we will specifically expand on three key directions: preprocessing & feature extraction, enhancement & noise reduction, and sound source separation.From a product perspective, this layer is almost "invisible" behind every audio product: real-time noise reduction in meeting software, post-production tuning for podcasts and short videos, the "voice enhancement mode" in voice recorders and phones, the "beautify voice" toggle on live streaming platforms, and the front-end preprocessing for ASR/speaker verification models — all of these are direct manifestations of waveform-level audio processing. Below, we continue to organize this from three angles — scenarios, principles, and models — and in subsequent subsections we will specifically expand on three key directions: preprocessing & feature extraction, enhancement & noise reduction, and sound source separation.
Waveform-level processing usually does not directly understand semantics; instead, it performs signal optimization around spectral structure and statistical characteristics:Waveform-level processing usually does not directly understand semantics; instead, it performs signal optimization around spectral structure and statistical characteristics:
Models at the waveform/spectrum level can be roughly divided into two categories: spectral-domain models and time-domain end-to-end models:Models at the waveform/spectrum level can be roughly divided into two categories: spectral-domain models and time-domain end-to-end models:
Any subsequent ASR, speaker verification, event detection, TTS, and other models require audio input that is as unified, clean, and structured as possible — this is the responsibility of the preprocessing and feature extraction layer. It handles the most basic yet critically important tasks of "clearing the stage" and "format unification," setting the stage for upstream audio models.Any subsequent ASR, speaker verification, event detection, TTS, and other models require audio input that is as unified, clean, and structured as possible — this is the responsibility of the preprocessing and feature extraction layer. It handles the most basic yet critically important tasks of "clearing the stage" and "format unification," setting the stage for upstream audio models.
In the preprocessing stage, the collected audio first undergoes sample rate conversion and channel conversion: for example, converting 48kHz stereo to 16kHz mono to meet the input specifications of downstream models and reduce computational cost. Subsequently, loudness normalization, DC offset removal, simple filtering, etc., are performed to make audio recorded from different devices and scenarios more consistent on the energy scale.In the preprocessing stage, the collected audio first undergoes sample rate conversion and channel conversion: for example, converting 48kHz stereo to 16kHz mono to meet the input specifications of downstream models and reduce computational cost. Subsequently, loudness normalization, DC offset removal, simple filtering, etc., are performed to make audio recorded from different devices and scenarios more consistent on the energy scale.
Voice Activity Detection (VAD) is another key component of preprocessing. It attempts to automatically segment "speech segments" and "silence/pure noise segments" in the audio stream, often based on frame energy, spectral entropy, zero-crossing rate, or small neural network discrimination. The benefit of VAD is that it can significantly reduce the amount of invalid data fed into ASR/speaker verification models, lowering computational load while preventing silence segments from interfering with recognition (e.g., being misrecognized as long strings of spaces or strange characters). In real-time communication, VAD can also drive the "voice activity indicator" and automatic mute logic.Voice Activity Detection (VAD) is another key component of preprocessing. It attempts to automatically segment "speech segments" and "silence/pure noise segments" in the audio stream, often based on frame energy, spectral entropy, zero-crossing rate, or small neural network discrimination. The benefit of VAD is that it can significantly reduce the amount of invalid data fed into ASR/speaker verification models, lowering computational load while preventing silence segments from interfering with recognition (e.g., being misrecognized as long strings of spaces or strange characters). In real-time communication, VAD can also drive the "voice activity indicator" and automatic mute logic.
At the feature extraction level, the most common approach is to convert the time-domain waveform into a spectrum or Mel-spectrogram. Through the Short-Time Fourier Transform (STFT), audio is decomposed into a frequency distribution that varies over time; through Mel filter banks, more perceptually relevant Mel-spectrograms or Mel cepstral features (such as log Mel-spectrogram, MFCC) can be obtained. These time–frequency features provide a "two-dimensional representation" for subsequent recognition, separation, and generation, analogous to grayscale images or multi-channel feature maps in vision, making them convenient for convolution, attention, and other architectures to process. With the development of end-to-end modeling, more and more models are learning features directly from waveforms (e.g., Wav2Vec 2.0), but in engineering practice, the combination of STFT + Mel features remains the most common and reliable front-end.At the feature extraction level, the most common approach is to convert the time-domain waveform into a spectrum or Mel-spectrogram. Through the Short-Time Fourier Transform (STFT), audio is decomposed into a frequency distribution that varies over time; through Mel filter banks, more perceptually relevant Mel-spectrograms or Mel cepstral features (such as log Mel-spectrogram, MFCC) can be obtained. These time–frequency features provide a "two-dimensional representation" for subsequent recognition, separation, and generation, analogous to grayscale images or multi-channel feature maps in vision, making them convenient for convolution, attention, and other architectures to process. With the development of end-to-end modeling, more and more models are learning features directly from waveforms (e.g., Wav2Vec 2.0), but in engineering practice, the combination of STFT + Mel features remains the most common and reliable front-end.
In real environments, sound almost always propagates through noise and reverberation: air conditioning hum, keyboard clicks, road noise, crowd chatter, and room echo all degrade the intelligibility and subjective quality of speech and music to varying degrees. The goal of speech enhancement and noise reduction is to suppress these background interferences while preserving the naturalness and completeness of speech as much as possible, restoring "muddy" sound to "clean" sound.In real environments, sound almost always propagates through noise and reverberation: air conditioning hum, keyboard clicks, road noise, crowd chatter, and room echo all degrade the intelligibility and subjective quality of speech and music to varying degrees. The goal of speech enhancement and noise reduction is to suppress these background interferences while preserving the naturalness and completeness of speech as much as possible, restoring "muddy" sound to "clean" sound.
In traditional methods, this task is primarily achieved through frequency-domain techniques such as spectral subtraction and Wiener filtering: first estimate the noise spectrum, then "subtract" the noise or adjust frequency band gains on the spectrum according to certain rules. While simple to implement and offering good real-time performance, these methods tend to produce noticeable "musical noise" and artifacts in scenarios with strong noise, non-stationary noise, and complex reverberation.In traditional methods, this task is primarily achieved through frequency-domain techniques such as spectral subtraction and Wiener filtering: first estimate the noise spectrum, then "subtract" the noise or adjust frequency band gains on the spectrum according to certain rules. While simple to implement and offering good real-time performance, these methods tend to produce noticeable "musical noise" and artifacts in scenarios with strong noise, non-stationary noise, and complex reverberation.
Deep learning methods, on the other hand, learn a mapping on the spectrum or waveform: given noisy speech, predict a time–frequency mask or directly predict the clean waveform. Common approaches include using Spectrogram-based U-Net, DCCRN, and other encoder–decoder architectures on Mel/linear spectrograms to carefully repair the spectrum of each frame; there are also end-to-end waveform enhancement methods using models such as Conv-TasNet, Demucs, Wave-U-Net directly on time-domain waveforms. These methods can significantly improve speech clarity and subjective listening quality in scenarios such as voice calls, online meetings, and recording restoration.Deep learning methods, on the other hand, learn a mapping on the spectrum or waveform: given noisy speech, predict a time–frequency mask or directly predict the clean waveform. Common approaches include using Spectrogram-based U-Net, DCCRN, and other encoder–decoder architectures on Mel/linear spectrograms to carefully repair the spectrum of each frame; there are also end-to-end waveform enhancement methods using models such as Conv-TasNet, Demucs, Wave-U-Net directly on time-domain waveforms. These methods can significantly improve speech clarity and subjective listening quality in scenarios such as voice calls, online meetings, and recording restoration.
In content creation and post-production, "recording restoration" often also involves reducing plosives, cutting sibilance, compensating for frequency band gaps, as well as equalization (EQ) and dynamic processing (compressor/limiter) — operations that are more "audio engineer" in nature. An increasing number of tools combine these traditional processes with deep models to provide one-click "audio repair" and "audio beautification" capabilities, serving podcasters, video creators, and live streaming platforms.In content creation and post-production, "recording restoration" often also involves reducing plosives, cutting sibilance, compensating for frequency band gaps, as well as equalization (EQ) and dynamic processing (compressor/limiter) — operations that are more "audio engineer" in nature. An increasing number of tools combine these traditional processes with deep models to provide one-click "audio repair" and "audio beautification" capabilities, serving podcasters, video creators, and live streaming platforms.
If enhancement and noise reduction are about "making the main sound more prominent and the background quieter," then sound source separation goes a step further by attempting to completely split multiple mixed sound sources into independent tracks. For example: multiple speakers talking simultaneously in a meeting recording; vocals and accompaniment mixed together in music; key events (such as alarms, shouting) buried in background noise in environmental recordings. The goal of sound source separation is to recover the waveform or spectrum of each individual sound source from a single or multiple mixed signals.If enhancement and noise reduction are about "making the main sound more prominent and the background quieter," then sound source separation goes a step further by attempting to completely split multiple mixed sound sources into independent tracks. For example: multiple speakers talking simultaneously in a meeting recording; vocals and accompaniment mixed together in music; key events (such as alarms, shouting) buried in background noise in environmental recordings. The goal of sound source separation is to recover the waveform or spectrum of each individual sound source from a single or multiple mixed signals.
In the speech domain, multi-speaker separation is a core application: the model needs to separate multiple overlapping voices into different channels based on speaker embeddings, time–frequency structure, and speaker characteristics, without separate microphone tracks. This capability not only improves the performance of multi-speaker ASR but also provides cleaner input for speaker diarization. In the music domain, vocals/accompaniment separation (singing voice separation) can extract clear vocal tracks and pure accompaniment tracks from a mixed song, used for covers, remixes, karaoke, music analysis, and more. Similarly, ambient sound/foreground sound separation can be used in security and IoT scenarios to extract key event sounds (such as glass breaking, conflict sounds) from complex backgrounds.In the speech domain, multi-speaker separation is a core application: the model needs to separate multiple overlapping voices into different channels based on speaker embeddings, time–frequency structure, and speaker characteristics, without separate microphone tracks. This capability not only improves the performance of multi-speaker ASR but also provides cleaner input for speaker diarization. In the music domain, vocals/accompaniment separation (singing voice separation) can extract clear vocal tracks and pure accompaniment tracks from a mixed song, used for covers, remixes, karaoke, music analysis, and more. Similarly, ambient sound/foreground sound separation can be used in security and IoT scenarios to extract key event sounds (such as glass breaking, conflict sounds) from complex backgrounds.
At the model level, sound source separation typically employs stronger modeling capabilities and more complex architectures than ordinary enhancement. End-to-end networks such as Conv-TasNet, Demucs, Wave-U-Net can directly perform multi-source decomposition in the time domain; in the spectral domain, multi-branch U-Net, attention, mask estimation, and other architectures are common, predicting specialized masks or spectra for different sound sources. With the growth of training data and computational resources, modern sound source separation models can already output high-quality separated tracks usable for practical creation and analysis in fairly complex reverberation and noise environments, providing a solid foundation for live voice beautification, multi-speaker meetings, music production, and audio retrieval.At the model level, sound source separation typically employs stronger modeling capabilities and more complex architectures than ordinary enhancement. End-to-end networks such as Conv-TasNet, Demucs, Wave-U-Net can directly perform multi-source decomposition in the time domain; in the spectral domain, multi-branch U-Net, attention, mask estimation, and other architectures are common, predicting specialized masks or spectra for different sound sources. With the growth of training data and computational resources, modern sound source separation models can already output high-quality separated tracks usable for practical creation and analysis in fairly complex reverberation and noise environments, providing a solid foundation for live voice beautification, multi-speaker meetings, music production, and audio retrieval.
After completing preprocessing, enhancement, and separation at the waveform level, we can finally begin to ask higher-level questions: "What is being said in the audio?" "Who is speaking?" "When is who speaking?" This layer focuses on various "understanding and annotation" tasks centered around speech itself: Automatic Speech Recognition (ASR), speaker recognition and verification, speaker diarization, and hotword and keyword detection (KWS) for interaction.After completing preprocessing, enhancement, and separation at the waveform level, we can finally begin to ask higher-level questions: "What is being said in the audio?" "Who is speaking?" "When is who speaking?" This layer focuses on various "understanding and annotation" tasks centered around speech itself: Automatic Speech Recognition (ASR), speaker recognition and verification, speaker diarization, and hotword and keyword detection (KWS) for interaction.
From a product perspective, this layer is the core of the vast majority of "voice products": voice input methods, meeting transcription, customer service recording analysis, intelligent customer service quality inspection, smart speaker and in-vehicle voice interaction, phone robots, and voiceprint verification in financial scenarios — almost all directly depend on these technologies. They transform the "clean sound" from the previous layer into text sequences, speaker labels, or keyword events, serving as one of the most important bridges from audio to the semantic world.From a product perspective, this layer is the core of the vast majority of "voice products": voice input methods, meeting transcription, customer service recording analysis, intelligent customer service quality inspection, smart speaker and in-vehicle voice interaction, phone robots, and voiceprint verification in financial scenarios — almost all directly depend on these technologies. They transform the "clean sound" from the previous layer into text sequences, speaker labels, or keyword events, serving as one of the most important bridges from audio to the semantic world.
Most tasks in this layer can be uniformly viewed as performing time alignment and sequence labeling on audio sequences:Most tasks in this layer can be uniformly viewed as performing time alignment and sequence labeling on audio sequences:
The model landscape for ASR and speaker technology includes both end-to-end architectures and specialized embedding models and clustering methods:The model landscape for ASR and speaker technology includes both end-to-end architectures and specialized embedding models and clustering methods:
Automatic Speech Recognition (ASR) is the main "audio → text" pathway: whether it's voice input methods, meeting transcription, smart captions, or customer service recording analysis, the first step is always to accurately convert what the user says into text. Modern ASR systems mostly adopt end-to-end architectures: starting from acoustic features (such as Mel-spectrograms or raw waveforms), going through a series of deep networks (such as Conformer, Citrinet, Transformer-based Encoders), and directly outputting text sequences or corresponding token sequences.Automatic Speech Recognition (ASR) is the main "audio → text" pathway: whether it's voice input methods, meeting transcription, smart captions, or customer service recording analysis, the first step is always to accurately convert what the user says into text. Modern ASR systems mostly adopt end-to-end architectures: starting from acoustic features (such as Mel-spectrograms or raw waveforms), going through a series of deep networks (such as Conformer, Citrinet, Transformer-based Encoders), and directly outputting text sequences or corresponding token sequences.
In terms of modeling, the main challenges of ASR include long-range dependencies, multilingualism and dialects, accent variations, overlapping speech, background noise, and domain-specific terminology. To address these, the current mainstream direction is to use large-scale unlabeled audio for self-supervised pre-training (such as Wav2Vec 2.0, HuBERT), or to perform large-scale supervised training on multilingual and multi-task data (such as Whisper), and then fine-tune with relatively small amounts of domain-specific data, thereby achieving good robustness across different languages, accents, and scenarios.In terms of modeling, the main challenges of ASR include long-range dependencies, multilingualism and dialects, accent variations, overlapping speech, background noise, and domain-specific terminology. To address these, the current mainstream direction is to use large-scale unlabeled audio for self-supervised pre-training (such as Wav2Vec 2.0, HuBERT), or to perform large-scale supervised training on multilingual and multi-task data (such as Whisper), and then fine-tune with relatively small amounts of domain-specific data, thereby achieving good robustness across different languages, accents, and scenarios.
At the product level, ASR is typically packaged as capabilities such as "voice input SDK," "cloud speech recognition API," and "meeting transcription service": the front-end can be real-time streaming recognition (RNN-T, streaming Transformer, etc.), while the back-end can enhance recognition of specific person names, place names, brand names, and business terminology through hotword injection, custom vocabularies, and contextual constraints. These recognition results often serve as the foundation for subsequent NLP, dialogue systems, and data analysis.At the product level, ASR is typically packaged as capabilities such as "voice input SDK," "cloud speech recognition API," and "meeting transcription service": the front-end can be real-time streaming recognition (RNN-T, streaming Transformer, etc.), while the back-end can enhance recognition of specific person names, place names, brand names, and business terminology through hotword injection, custom vocabularies, and contextual constraints. These recognition results often serve as the foundation for subsequent NLP, dialogue systems, and data analysis.
Compared to "what is being said," "who is speaking" is equally important in many applications: scenarios such as finance, government, customer service, and security require voiceprint recognition to verify identity or investigate risks; while meeting and interview scenarios need to know "who said each sentence" to support speaker-attributed transcription, speaking statistics, and behavioral analysis.Compared to "what is being said," "who is speaking" is equally important in many applications: scenarios such as finance, government, customer service, and security require voiceprint recognition to verify identity or investigate risks; while meeting and interview scenarios need to know "who said each sentence" to support speaker-attributed transcription, speaking statistics, and behavioral analysis.
In the Speaker Recognition/Verification task, the system's goal is: given a segment of speech, determine who the speaker is, or determine whether it matches a registered speaker. Modern systems typically use models such as ECAPA-TDNN and x-vector to extract a fixed-dimensional speaker embedding vector from the speech segment. During the training phase, a combination of speaker classification and metric learning ensures that embeddings of the same person are more clustered and the distance between embeddings of different people is larger; during the inference phase, nearest neighbor or back-end discriminators (such as PLDA, Cosine scoring with margin) are used for verification and recognition. In this way, the system can answer "is it the same person" with a certain confidence level across phone, microphone, and noisy environments.In the Speaker Recognition/Verification task, the system's goal is: given a segment of speech, determine who the speaker is, or determine whether it matches a registered speaker. Modern systems typically use models such as ECAPA-TDNN and x-vector to extract a fixed-dimensional speaker embedding vector from the speech segment. During the training phase, a combination of speaker classification and metric learning ensures that embeddings of the same person are more clustered and the distance between embeddings of different people is larger; during the inference phase, nearest neighbor or back-end discriminators (such as PLDA, Cosine scoring with margin) are used for verification and recognition. In this way, the system can answer "is it the same person" with a certain confidence level across phone, microphone, and noisy environments.
Speaker Diarization further answers "who is speaking when." The traditional approach typically involves three steps: first use VAD to find segments with speech, then cut the long audio into short segments, extract speaker embeddings for each segment, and finally perform clustering and temporal stitching in the embedding space to obtain a multi-speaker timeline. More advanced End-to-End Diarization (EEND) methods attempt to directly output a "time × speaker" boolean matrix from audio features, learning complex patterns such as overlapping speech and speaker changes in an end-to-end manner. Diarization is highly valuable in scenarios such as meetings, interview programs, court records, and phone customer service, often combined with ASR to produce "text records with speaker labels."Speaker Diarization further answers "who is speaking when." The traditional approach typically involves three steps: first use VAD to find segments with speech, then cut the long audio into short segments, extract speaker embeddings for each segment, and finally perform clustering and temporal stitching in the embedding space to obtain a multi-speaker timeline. More advanced End-to-End Diarization (EEND) methods attempt to directly output a "time × speaker" boolean matrix from audio features, learning complex patterns such as overlapping speech and speaker changes in an end-to-end manner. Diarization is highly valuable in scenarios such as meetings, interview programs, court records, and phone customer service, often combined with ASR to produce "text records with speaker labels."
In a continuous audio stream, not every second is worth being fully recognized and stored. The role of hotword and keyword detection (KWS) is that of an always-on "gatekeeper":In a continuous audio stream, not every second is worth being fully recognized and stored. The role of hotword and keyword detection (KWS) is that of an always-on "gatekeeper":
In terms of technical implementation, KWS typically needs to operate under constraints of extremely low computational cost and low latency, especially for wake word detection on local devices: the model is often a small CNN/RNN/Transformer front-end connected to a CTC or gated discrimination head, detecting acoustic patterns of specific words, and using sliding windows and confidence smoothing to avoid false wakes. For keyword quality inspection scenarios, stronger ASR + keyword matching/regex + statistical analysis can be used, or end-to-end keyword tagging models can be trained directly. Regardless of the form, KWS essentially adds a layer of "event-level" semantic filtering on the speech stream, serving as an important interface connecting the audio world with interaction logic.In terms of technical implementation, KWS typically needs to operate under constraints of extremely low computational cost and low latency, especially for wake word detection on local devices: the model is often a small CNN/RNN/Transformer front-end connected to a CTC or gated discrimination head, detecting acoustic patterns of specific words, and using sliding windows and confidence smoothing to avoid false wakes. For keyword quality inspection scenarios, stronger ASR + keyword matching/regex + statistical analysis can be used, or end-to-end keyword tagging models can be trained directly. Regardless of the form, KWS essentially adds a layer of "event-level" semantic filtering on the speech stream, serving as an important interface connecting the audio world with interaction logic.
Not all audio is centered on "speech." In reality, there are many scenarios related to environmental sounds, event sounds, and music, which are more concerned with: "What sound event occurred?" "What is the current acoustic scene?" "What style is this song, what instruments are used, what are the rhythm and key?" This set of capabilities is collectively referred to as audio/music understanding, primarily revolving around sound event detection, environmental/scene classification, and music attribute understanding.Not all audio is centered on "speech." In reality, there are many scenarios related to environmental sounds, event sounds, and music, which are more concerned with: "What sound event occurred?" "What is the current acoustic scene?" "What style is this song, what instruments are used, what are the rhythm and key?" This set of capabilities is collectively referred to as audio/music understanding, primarily revolving around sound event detection, environmental/scene classification, and music attribute understanding.
From a product perspective, audio understanding technology supports a wide range of applications such as acoustic security monitoring, IoT acoustic sensors, environmental adaptation for smart devices, music recommendation and classification, music copyright identification, music retrieval, and creative assistance. Similar to "image classification + fine-grained classification" in vision, this layer structures the originally continuous and complex sound space into discrete event labels, multi-dimensional attribute vectors, and style descriptions.From a product perspective, audio understanding technology supports a wide range of applications such as acoustic security monitoring, IoT acoustic sensors, environmental adaptation for smart devices, music recommendation and classification, music copyright identification, music retrieval, and creative assistance. Similar to "image classification + fine-grained classification" in vision, this layer structures the originally continuous and complex sound space into discrete event labels, multi-dimensional attribute vectors, and style descriptions.
Audio/music understanding is mostly based on time–frequency features + deep neural networks for classification or multi-label annotation:Audio/music understanding is mostly based on time–frequency features + deep neural networks for classification or multi-label annotation:
Common audio understanding models are mostly pre-trained on public datasets (such as AudioSet) and then transferred to specific tasks:Common audio understanding models are mostly pre-trained on public datasets (such as AudioSet) and then transferred to specific tasks:
In security, IoT, smart cities, and in-vehicle systems, cameras alone are not sufficient to fully understand environmental states. The goal of sound event detection is to enable systems to "hear and understand" key events: when glass breaks, an alarm sounds, a baby cries, a collision, scream, fight, or destructive behavior occurs, the system can identify and alert from the audio signal. Unlike speech recognition, such events are often short, non-verbal, with varying frequency ranges and energy patterns, and may be highly overlapped with background noise.In security, IoT, smart cities, and in-vehicle systems, cameras alone are not sufficient to fully understand environmental states. The goal of sound event detection is to enable systems to "hear and understand" key events: when glass breaks, an alarm sounds, a baby cries, a collision, scream, fight, or destructive behavior occurs, the system can identify and alert from the audio signal. Unlike speech recognition, such events are often short, non-verbal, with varying frequency ranges and energy patterns, and may be highly overlapped with background noise.
Environmental/scene classification focuses more on persistent acoustic scenes: is it a quiet office, a bustling street, inside a car, a high-speed rail station, or a café? The system can automatically adjust noise reduction intensity, echo cancellation parameters, microphone array beam steering based on the acoustic scene, and even change interaction strategies (e.g., using shorter feedback interactions in a car, increasing output volume on noisy streets). In IoT scenarios, an "acoustic network" composed of multiple sound sensors can be used for long-term monitoring and statistical analysis of environmental states.Environmental/scene classification focuses more on persistent acoustic scenes: is it a quiet office, a bustling street, inside a car, a high-speed rail station, or a café? The system can automatically adjust noise reduction intensity, echo cancellation parameters, microphone array beam steering based on the acoustic scene, and even change interaction strategies (e.g., using shorter feedback interactions in a car, increasing output volume on noisy streets). In IoT scenarios, an "acoustic network" composed of multiple sound sensors can be used for long-term monitoring and statistical analysis of environmental states.
In terms of technical implementation, both types of tasks mostly adopt multi-label classification + temporal modeling approaches: convert audio to Mel-spectrograms, use VGGish, PANNs, AST, or similar models for feature extraction, and then use temporal pooling or sequence models to output the activation of each label on the time axis. Since many datasets only provide "clip-level labels" (weak labels), models often need to learn the temporal localization of events under weak supervision through methods such as multi-instance learning and self-attention pooling.In terms of technical implementation, both types of tasks mostly adopt multi-label classification + temporal modeling approaches: convert audio to Mel-spectrograms, use VGGish, PANNs, AST, or similar models for feature extraction, and then use temporal pooling or sequence models to output the activation of each label on the time axis. Since many datasets only provide "clip-level labels" (weak labels), models often need to learn the temporal localization of events under weak supervision through methods such as multi-instance learning and self-attention pooling.
In the music domain, the goal of audio understanding is not merely "what song is this," but also to answer: "What style is this song? What instruments are used? How fast or slow is the tempo? What is the key and general harmonic structure?" This information supports music recommendation and playlist curation on the one hand, and on the other hand provides structured "music metadata" for creators and generative models.In the music domain, the goal of audio understanding is not merely "what song is this," but also to answer: "What style is this song? What instruments are used? How fast or slow is the tempo? What is the key and general harmonic structure?" This information supports music recommendation and playlist curation on the one hand, and on the other hand provides structured "music metadata" for creators and generative models.
The genre classification task categorizes songs into different styles such as pop, rock, classical, hip-hop, electronic, Lo-Fi based on their overall acoustic features and structure; instrument recognition distinguishes the acoustic fingerprints of different instruments such as drums, bass, guitar, piano, and strings on time–frequency features, which can be used for instrument statistics, music retrieval, and mixing analysis. Tempo/key analysis estimates BPM, beat positions, time signature, and key, providing a foundation for tasks such as beat matching, automatic harmonization, DJ mixing, and game soundtrack synchronization.The genre classification task categorizes songs into different styles such as pop, rock, classical, hip-hop, electronic, Lo-Fi based on their overall acoustic features and structure; instrument recognition distinguishes the acoustic fingerprints of different instruments such as drums, bass, guitar, piano, and strings on time–frequency features, which can be used for instrument statistics, music retrieval, and mixing analysis. Tempo/key analysis estimates BPM, beat positions, time signature, and key, providing a foundation for tasks such as beat matching, automatic harmonization, DJ mixing, and game soundtrack synchronization.
In terms of models, music understanding largely follows general audio models (such as PANNs, AST), but there are also many models and pre-trained embeddings specifically oriented toward Music Information Retrieval (MIR). A typical approach is to perform multi-label music tag learning (genre, mood, instrument, era, etc.) on large-scale music datasets to obtain a music embedding space, and then fine-tune or perform zero-shot inference on the specific tasks mentioned above. By combining these models, music platforms can more intelligently complete music classification and recommendation, copyright platforms can enhance music fingerprinting and similarity retrieval, and creative tools can leverage these understanding capabilities to recommend suitable accompaniment to users, extend similar styles, or automatically generate musical structures.In terms of models, music understanding largely follows general audio models (such as PANNs, AST), but there are also many models and pre-trained embeddings specifically oriented toward Music Information Retrieval (MIR). A typical approach is to perform multi-label music tag learning (genre, mood, instrument, era, etc.) on large-scale music datasets to obtain a music embedding space, and then fine-tune or perform zero-shot inference on the specific tasks mentioned above. By combining these models, music platforms can more intelligently complete music classification and recommendation, copyright platforms can enhance music fingerprinting and similarity retrieval, and creative tools can leverage these understanding capabilities to recommend suitable accompaniment to users, extend similar styles, or automatically generate musical structures.
After completing the "cleanup," "recognition," and "understanding" of audio, the natural next-layer question is: "Can we directly make machines 'speak,' 'sing,' or even 'compose'?" This is the world of speech and audio generation: from Text-to-Speech (TTS), from one voice to another (VC / Voice Cloning), to broader music and sound effect generation, and further to singing voice synthesis that can sing lyrics and melodies. Similar to image generation, this layer is no longer just about labeling or extracting structure from existing data but actively "creating" new sound content.After completing the "cleanup," "recognition," and "understanding" of audio, the natural next-layer question is: "Can we directly make machines 'speak,' 'sing,' or even 'compose'?" This is the world of speech and audio generation: from Text-to-Speech (TTS), from one voice to another (VC / Voice Cloning), to broader music and sound effect generation, and further to singing voice synthesis that can sing lyrics and melodies. Similar to image generation, this layer is no longer just about labeling or extracting structure from existing data but actively "creating" new sound content.
At the product level, this layer's capabilities have already permeated various applications: voice product lines such as OpenAI TTS, ElevenLabs, ByteDance Volcano Engine, and MiniMax provide high-quality synthesized speech for applications; music generation platforms such as Suno and Udio provide creators and even ordinary users with the ability to go from text prompts to complete music; games, videos, virtual streamers, and digital humans rely on these models for dubbing and singing, greatly lowering the barrier to content production.At the product level, this layer's capabilities have already permeated various applications: voice product lines such as OpenAI TTS, ElevenLabs, ByteDance Volcano Engine, and MiniMax provide high-quality synthesized speech for applications; music generation platforms such as Suno and Udio provide creators and even ordinary users with the ability to go from text prompts to complete music; games, videos, virtual streamers, and digital humans rely on these models for dubbing and singing, greatly lowering the barrier to content production.
Speech and audio generation typically adopts a layered modeling approach of "high-level representation → low-level waveform":Speech and audio generation typically adopts a layered modeling approach of "high-level representation → low-level waveform":
The current mainstream technical approaches for speech and audio generation include:The current mainstream technical approaches for speech and audio generation include:
Text-to-Speech (TTS) is the most intuitive speech generation task: input a piece of text, output a segment of natural and fluent speech that ideally is nearly indistinguishable from a human voice. Modern TTS systems are typically divided into two main stages: text to acoustic features (such as Mel-spectrogram), and acoustic features to waveform.Text-to-Speech (TTS) is the most intuitive speech generation task: input a piece of text, output a segment of natural and fluent speech that ideally is nearly indistinguishable from a human voice. Modern TTS systems are typically divided into two main stages: text to acoustic features (such as Mel-spectrogram), and acoustic features to waveform.
In the first stage, the model needs to handle problems such as tokenization, phonemization, polyphone disambiguation, punctuation and pausing, and prosody prediction. Typical models include the attention-based Tacotron series and the length-prediction-based FastSpeech series, with the latter significantly accelerating synthesis and improving stability through a non-autoregressive architecture. In recent years, end-to-end models such as VITS have merged acoustic modeling and the vocoder into a unified framework, further simplifying the system.In the first stage, the model needs to handle problems such as tokenization, phonemization, polyphone disambiguation, punctuation and pausing, and prosody prediction. Typical models include the attention-based Tacotron series and the length-prediction-based FastSpeech series, with the latter significantly accelerating synthesis and improving stability through a non-autoregressive architecture. In recent years, end-to-end models such as VITS have merged acoustic modeling and the vocoder into a unified framework, further simplifying the system.
In the second stage, Neural Vocoders such as WaveNet, WaveRNN, HiFi-GAN, WaveGlow, etc., are responsible for converting Mel-spectrograms or other intermediate representations into high-fidelity waveforms. A well-trained vocoder can not only generate natural and clear speech but also faithfully reproduce different timbres, emotions, and styles. Modern TTS systems also support multi-speaker modeling (via speaker embedding), timbre/speed/emotion control (such as "excited," "calm," "broadcasting style"), and cross-lingual TTS, providing highly customizable voice capabilities for various applications.In the second stage, Neural Vocoders such as WaveNet, WaveRNN, HiFi-GAN, WaveGlow, etc., are responsible for converting Mel-spectrograms or other intermediate representations into high-fidelity waveforms. A well-trained vocoder can not only generate natural and clear speech but also faithfully reproduce different timbres, emotions, and styles. Modern TTS systems also support multi-speaker modeling (via speaker embedding), timbre/speed/emotion control (such as "excited," "calm," "broadcasting style"), and cross-lingual TTS, providing highly customizable voice capabilities for various applications.
In many creative and assistive scenarios, we want to change the speaker's timbre or style without altering the content and prosody — this is the task of Voice Conversion (VC) and Voice Cloning. The former primarily addresses "turning person A's speech into person B's voice"; the latter further emphasizes "learning a new timbre from just a few samples or even a few seconds of speech."In many creative and assistive scenarios, we want to change the speaker's timbre or style without altering the content and prosody — this is the task of Voice Conversion (VC) and Voice Cloning. The former primarily addresses "turning person A's speech into person B's voice"; the latter further emphasizes "learning a new timbre from just a few samples or even a few seconds of speech."
Technically, VC typically adopts a "content–timbre decoupling" approach: a content encoder extracts speech content and prosody information (which can be discrete units based on ASR or continuous representations from self-supervised learning), and then a conditional generator combines the target speaker embedding or codec conditions to generate new speech with the target timbre but largely unchanged semantics and rhythm. With the introduction of neural codecs, speech can be directly edited in the codec space to achieve high-fidelity conversion.Technically, VC typically adopts a "content–timbre decoupling" approach: a content encoder extracts speech content and prosody information (which can be discrete units based on ASR or continuous representations from self-supervised learning), and then a conditional generator combines the target speaker embedding or codec conditions to generate new speech with the target timbre but largely unchanged semantics and rhythm. With the introduction of neural codecs, speech can be directly edited in the codec space to achieve high-fidelity conversion.
Voice Cloning builds on VC with an emphasis on few-shot and generalization capabilities: the model needs to extract stable speaker representations from a few samples or even a few seconds of audio, and based on this, generate synthesized speech with consistent style and similar timbre. This capability is very useful in virtual personas, personalized assistants, game character customization, dubbing acceleration, and more, but it must also strictly comply with legal and ethical norms, ensuring use only under conditions of compliant authorization, full informed consent, and security controls, to avoid risks of misuse or identity impersonation.Voice Cloning builds on VC with an emphasis on few-shot and generalization capabilities: the model needs to extract stable speaker representations from a few samples or even a few seconds of audio, and based on this, generate synthesized speech with consistent style and similar timbre. This capability is very useful in virtual personas, personalized assistants, game character customization, dubbing acceleration, and more, but it must also strictly comply with legal and ethical norms, ensuring use only under conditions of compliant authorization, full informed consent, and security controls, to avoid risks of misuse or identity impersonation.
Compared to speech generation, music and sound effect generation is more complex in structure and time scale: music often lasts longer with richer internal structure (sections, melody, harmony, rhythm); sound effects are diverse, ranging from natural environments (rain, wind, ocean waves) to foley sounds (UI clicks, notification sounds, game skill effects), each with their own patterns. In recent years, models based on neural codecs, sequence modeling, and diffusion have made "generating complete music/sound effects from text" a reality.Compared to speech generation, music and sound effect generation is more complex in structure and time scale: music often lasts longer with richer internal structure (sections, melody, harmony, rhythm); sound effects are diverse, ranging from natural environments (rain, wind, ocean waves) to foley sounds (UI clicks, notification sounds, game skill effects), each with their own patterns. In recent years, models based on neural codecs, sequence modeling, and diffusion have made "generating complete music/sound effects from text" a reality.
In music generation, models like MusicLM, MusicGen, Suno, and Udio typically encode audio into discrete codec token sequences and then train text-conditioned or multimodal-conditioned generative models in this discrete space. Users only need to provide a text description (such as "moderate tempo, warm and healing Lo-Fi background music, suitable for studying and focusing," "tense electronic orchestral score, suitable for a sci-fi trailer"), or upload a reference music clip, and the model can generate high-quality music lasting tens of seconds or even minutes. For creators, this is both a source of inspiration and a powerful tool for rapid prototyping and background music generation.In music generation, models like MusicLM, MusicGen, Suno, and Udio typically encode audio into discrete codec token sequences and then train text-conditioned or multimodal-conditioned generative models in this discrete space. Users only need to provide a text description (such as "moderate tempo, warm and healing Lo-Fi background music, suitable for studying and focusing," "tense electronic orchestral score, suitable for a sci-fi trailer"), or upload a reference music clip, and the model can generate high-quality music lasting tens of seconds or even minutes. For creators, this is both a source of inspiration and a powerful tool for rapid prototyping and background music generation.
In sound effect generation, similar techniques can generate UI sound effects, notification sounds, game ambient sounds, etc., from text prompts, helping product and game teams rapidly iterate on sound design. Combined with the audio understanding capabilities of the previous layer, style alignment and scene adaptation can also be achieved, such as automatically matching sound effect styles based on visuals or game levels.In sound effect generation, similar techniques can generate UI sound effects, notification sounds, game ambient sounds, etc., from text prompts, helping product and game teams rapidly iterate on sound design. Combined with the audio understanding capabilities of the previous layer, style alignment and scene adaptation can also be achieved, such as automatically matching sound effect styles based on visuals or game levels.
Whether it is speech or music and sound effect generation, this layer of capability is rapidly evolving: from the early days of heavily synthesized machine sounds to today's high-fidelity content that is nearly indistinguishable from human voices and professional music. At the same time, issues surrounding copyright, compliance, traceability, and controllability have become increasingly important — how to provide powerful creative tools while protecting the legitimate rights and interests of creators and users will remain a key topic that this layer of technology must continuously address.Whether it is speech or music and sound effect generation, this layer of capability is rapidly evolving: from the early days of heavily synthesized machine sounds to today's high-fidelity content that is nearly indistinguishable from human voices and professional music. At the same time, issues surrounding copyright, compliance, traceability, and controllability have become increasingly important — how to provide powerful creative tools while protecting the legitimate rights and interests of creators and users will remain a key topic that this layer of technology must continuously address.
In the multimodal AI system, the video modality is responsible for understanding and generating "visual signals that change over time." Compared to single-frame images, video not only contains spatial information such as texture, shape, and layout, but also carries rich temporal cues: the rise and fall of actions, the motion trajectories of objects, the rhythm of shot transitions, and so on. Whether it's behavior recognition in surveillance, motion analysis in sports training, one-click editing on short-video platforms, or intelligent parsing of long videos, they all fundamentally rely on a comprehensive set of understanding and generation capabilities built around "frame sequences."In the multimodal AI system, the video modality is responsible for understanding and generating "visual signals that change over time." Compared to single-frame images, video not only contains spatial information such as texture, shape, and layout, but also carries rich temporal cues: the rise and fall of actions, the motion trajectories of objects, the rhythm of shot transitions, and so on. Whether it's behavior recognition in surveillance, motion analysis in sports training, one-click editing on short-video platforms, or intelligent parsing of long videos, they all fundamentally rely on a comprehensive set of understanding and generation capabilities built around "frame sequences."
From an engineering perspective, video capabilities can be roughly divided into several layers: low-level video enhancement and restoration ensures that content is "clear enough to see"; video understanding and structural analysis answers the question of "what is happening"; building on that, video + language multimodal tasks convert video content into structured descriptions and retrieval interfaces usable by text; further, video generation and editing works in reverse, starting from text or example videos to generate or reassemble video content in a controllable manner; and a class of applications represented by digital humans / virtual avatars integrates speech, language, motion, and video rendering together, forming a new paradigm for interaction and content production.From an engineering perspective, video capabilities can be roughly divided into several layers: low-level video enhancement and restoration ensures that content is "clear enough to see"; video understanding and structural analysis answers the question of "what is happening"; building on that, video + language multimodal tasks convert video content into structured descriptions and retrieval interfaces usable by text; further, video generation and editing works in reverse, starting from text or example videos to generate or reassemble video content in a controllable manner; and a class of applications represented by digital humans / virtual avatars integrates speech, language, motion, and video rendering together, forming a new paradigm for interaction and content production.
Below, we similarly start from layered capabilities to organize video-related capabilities.Below, we similarly start from layered capabilities to organize video-related capabilities.
At the most fundamental level of video technology, the first concern is not "who is in the frame" or "what event is happening," but whether the video itself is stable, clear, and comfortable: whether the picture shakes, is blurry, has excessive noise, or whether the aspect ratio suits the target playback device. The traditional video processing layer primarily works at the level of frame sequences and spatiotemporal pixels, using operations such as enhancement, restoration, super-resolution, frame interpolation, and reframing to convert noisy, shaky, low-resolution, or improperly proportioned raw video into "high-quality temporal signals" that are more suitable for viewing and subsequent analysis. It can be analogized to "image restoration and enhancement + geometric correction" in the image modality, except that here, smoothness and consistency across the time dimension are additionally introduced.At the most fundamental level of video technology, the first concern is not "who is in the frame" or "what event is happening," but whether the video itself is stable, clear, and comfortable: whether the picture shakes, is blurry, has excessive noise, or whether the aspect ratio suits the target playback device. The traditional video processing layer primarily works at the level of frame sequences and spatiotemporal pixels, using operations such as enhancement, restoration, super-resolution, frame interpolation, and reframing to convert noisy, shaky, low-resolution, or improperly proportioned raw video into "high-quality temporal signals" that are more suitable for viewing and subsequent analysis. It can be analogized to "image restoration and enhancement + geometric correction" in the image modality, except that here, smoothness and consistency across the time dimension are additionally introduced.
From a product perspective, this layer of capabilities is almost "invisible" behind all video products: one-click quality enhancement in editing software, automatic quality upgrades on short-video platforms, intelligent super-resolution and frame interpolation in TV boxes and players, film restoration services, and multi-frame preprocessing for upstream detection/recognition models are all direct manifestations of traditional video processing. Below, we continue to organize from three angles — scenarios, principles, and models — and expand on key directions such as video enhancement and restoration, super-resolution, and frame interpolation in subsequent subsections.From a product perspective, this layer of capabilities is almost "invisible" behind all video products: one-click quality enhancement in editing software, automatic quality upgrades on short-video platforms, intelligent super-resolution and frame interpolation in TV boxes and players, film restoration services, and multi-frame preprocessing for upstream detection/recognition models are all direct manifestations of traditional video processing. Below, we continue to organize from three angles — scenarios, principles, and models — and expand on key directions such as video enhancement and restoration, super-resolution, and frame interpolation in subsequent subsections.
In online video platforms, editing tools, surveillance systems, and terminal devices, traditional video processing mainly appears in the following typical scenarios:In online video platforms, editing tools, surveillance systems, and terminal devices, traditional video processing mainly appears in the following typical scenarios:
Traditional video processing typically does not directly understand semantic categories; instead, it models and optimizes around quality, stability, and temporal consistency at the spatiotemporal signal level:Traditional video processing typically does not directly understand semantic categories; instead, it models and optimizes around quality, stability, and temporal consistency at the spatiotemporal signal level:
In concrete implementation, traditional video processing comprehensively uses classical video signal processing methods and deep learning models, striking a balance between effectiveness, efficiency, and deployment form:In concrete implementation, traditional video processing comprehensively uses classical video signal processing methods and deep learning models, striking a balance between effectiveness, efficiency, and deployment form:
Taken together, this layer primarily lays the physical and perceptual foundation for video "before semantics": it helps users obtain a more comfortable viewing experience and also provides cleaner, more stable input for upstream detection, recognition, and generation models. Below, we expand on the sub-directions of video enhancement and restoration and super-resolution and frame interpolation.Taken together, this layer primarily lays the physical and perceptual foundation for video "before semantics": it helps users obtain a more comfortable viewing experience and also provides cleaner, more stable input for upstream detection, recognition, and generation models. Below, we expand on the sub-directions of video enhancement and restoration and super-resolution and frame interpolation.
Under real shooting conditions, video is often not "clean": severe shaking from handheld devices, high noise and a smeared look in low light, block artifacts and color banding from network compression, and fading and scratches from old equipment all make video quality noticeably below the ideal. The goal of video enhancement and restoration is to restore a stable, clear, and natural viewing experience to the greatest extent possible without changing the semantic content of the video — polishing material that is "barely watchable" to a level that is "pleasing or even good-looking."Under real shooting conditions, video is often not "clean": severe shaking from handheld devices, high noise and a smeared look in low light, block artifacts and color banding from network compression, and fading and scratches from old equipment all make video quality noticeably below the ideal. The goal of video enhancement and restoration is to restore a stable, clear, and natural viewing experience to the greatest extent possible without changing the semantic content of the video — polishing material that is "barely watchable" to a level that is "pleasing or even good-looking."
In the temporal domain, enhancement and restoration must first address the problem of stability. By performing feature matching or optical flow estimation on consecutive frames, global camera motion and local object motion can be separated, and the smoothed camera trajectory is then used to re-render output frames, thereby suppressing rapid jitter and subtle shaking and preventing viewers from experiencing dizziness during viewing. On this foundation, frame-level denoising, deblurring, and artifact removal focus more on joint spatial–temporal modeling: multi-frame joint denoising uses redundant information from preceding and following frames, performing processing similar to "multi-exposure fusion" in the temporal direction, effectively suppressing high-ISO noise and compression noise while preserving detail texture; for mild motion blur, deconvolution-style sharpening is performed on the frame sequence by estimating the blur kernel or using an end-to-end deep network, making both static backgrounds and moving subjects sharper.In the temporal domain, enhancement and restoration must first address the problem of stability. By performing feature matching or optical flow estimation on consecutive frames, global camera motion and local object motion can be separated, and the smoothed camera trajectory is then used to re-render output frames, thereby suppressing rapid jitter and subtle shaking and preventing viewers from experiencing dizziness during viewing. On this foundation, frame-level denoising, deblurring, and artifact removal focus more on joint spatial–temporal modeling: multi-frame joint denoising uses redundant information from preceding and following frames, performing processing similar to "multi-exposure fusion" in the temporal direction, effectively suppressing high-ISO noise and compression noise while preserving detail texture; for mild motion blur, deconvolution-style sharpening is performed on the frame sequence by estimating the blur kernel or using an end-to-end deep network, making both static backgrounds and moving subjects sharper.
For old films and low-quality material, restoration also involves "reconstruction" at the color and structure level. Film aging causes yellowing, reduced contrast, and noticeable local scratches and blemishes; early digital video commonly suffers from low resolution, heavy compression, and edge aliasing. Modern restoration workflows often use multi-step collaboration: first, detection and segmentation models locate locally damaged areas such as scratches and blemishes; then, spatiotemporal inpainting networks "borrow material to fill holes" from neighboring frames and neighboring spatial pixels; simultaneously, color restoration and contrast reshaping bring the overall tone close to the original shooting or intended style reference. For heavily compressed video, dedicated artifact-removal networks targeting blocking artifacts and ringing artifacts are also introduced, improving edges and details without excessive smoothing.For old films and low-quality material, restoration also involves "reconstruction" at the color and structure level. Film aging causes yellowing, reduced contrast, and noticeable local scratches and blemishes; early digital video commonly suffers from low resolution, heavy compression, and edge aliasing. Modern restoration workflows often use multi-step collaboration: first, detection and segmentation models locate locally damaged areas such as scratches and blemishes; then, spatiotemporal inpainting networks "borrow material to fill holes" from neighboring frames and neighboring spatial pixels; simultaneously, color restoration and contrast reshaping bring the overall tone close to the original shooting or intended style reference. For heavily compressed video, dedicated artifact-removal networks targeting blocking artifacts and ringing artifacts are also introduced, improving edges and details without excessive smoothing.
These enhancement and restoration capabilities are often presented as "one-click" in products: the user simply checks "stabilization," "quality enhancement," or "old video restoration," and the system automatically selects the appropriate model and parameter combination in the background, performing multi-stage processing on the video frame sequence. For the business, this layer directly determines the audience's subjective evaluation of image quality and indirectly affects the performance of upstream analysis models: cleaner, more stable video input often means more reliable face/license plate recognition, more accurate behavior detection, and fewer false alarms.These enhancement and restoration capabilities are often presented as "one-click" in products: the user simply checks "stabilization," "quality enhancement," or "old video restoration," and the system automatically selects the appropriate model and parameter combination in the background, performing multi-stage processing on the video frame sequence. For the business, this layer directly determines the audience's subjective evaluation of image quality and indirectly affects the performance of upstream analysis models: cleaner, more stable video input often means more reliable face/license plate recognition, more accurate behavior detection, and fewer false alarms.
Against the backdrop of ever-upgrading display devices and increasing user demands for detail and smoothness, a large amount of existing video content appears "innately inadequate" in resolution and frame rate: 1080p looks insufficiently sharp on 4K screens, and 24/30fps tends to show ghosting or stuttering on large screens and in fast-motion scenes. Super-resolution and frame interpolation technologies are designed to solve these two problems: the former "fills in details" in the spatial dimension, and the latter "fills in the process" in the temporal dimension, together elevating video that is "barely clear" to a viewing experience that is "rich in detail and smoothly playing."Against the backdrop of ever-upgrading display devices and increasing user demands for detail and smoothness, a large amount of existing video content appears "innately inadequate" in resolution and frame rate: 1080p looks insufficiently sharp on 4K screens, and 24/30fps tends to show ghosting or stuttering on large screens and in fast-motion scenes. Super-resolution and frame interpolation technologies are designed to solve these two problems: the former "fills in details" in the spatial dimension, and the latter "fills in the process" in the temporal dimension, together elevating video that is "barely clear" to a viewing experience that is "rich in detail and smoothly playing."
Video super-resolution adds one key dimension compared to single-frame image super-resolution: time. Simple frame-by-frame upscaling tends to cause inconsistency in details between adjacent frames, resulting in flickering and texture jitter. Therefore, mainstream methods all leverage information from multiple surrounding frames, using optical flow estimation or feature-level alignment to align details from neighboring frames to the target frame, and then performing detail reconstruction after alignment. Models such as EDVR, BasicVSR / BasicVSR++, and the video version of Real‑ESRGAN first align and aggregate multiple frames in feature space, then use deep networks to infer high-resolution details, avoiding the "blur" and "plastic feel" that come with simple interpolation. In this process, balancing "physical plausibility" and "perceptual attractiveness" is the core of loss design and training strategy: both objective metrics (such as PSNR, SSIM) must be improved, and subjective viewing must feel natural, without over-sharpening or false details.Video super-resolution adds one key dimension compared to single-frame image super-resolution: time. Simple frame-by-frame upscaling tends to cause inconsistency in details between adjacent frames, resulting in flickering and texture jitter. Therefore, mainstream methods all leverage information from multiple surrounding frames, using optical flow estimation or feature-level alignment to align details from neighboring frames to the target frame, and then performing detail reconstruction after alignment. Models such as EDVR, BasicVSR / BasicVSR++, and the video version of Real‑ESRGAN first align and aggregate multiple frames in feature space, then use deep networks to infer high-resolution details, avoiding the "blur" and "plastic feel" that come with simple interpolation. In this process, balancing "physical plausibility" and "perceptual attractiveness" is the core of loss design and training strategy: both objective metrics (such as PSNR, SSIM) must be improved, and subjective viewing must feel natural, without over-sharpening or false details.
Frame interpolation focuses on "filling frames" along the time axis. Traditional methods rely on optical flow estimation, first predicting the motion of each pixel between two frames, then interpolating at an intermediate position according to certain rules to generate a new frame. However, in areas of fast motion, multi-object occlusion, or complex texture, optical flow is often not accurate enough, easily producing ghosting, double images, or local deformation. Deep frame interpolation models such as DAIN, RIFE, and FILM simultaneously learn fusion strategies for optical flow, depth, or intermediate features through end-to-end networks, directly outputting interpolated frames, with noticeably improved stability and visual quality in complex scenes. For sports events, action game recordings, and slow-motion creation, frame interpolation can smoothly boost 24/30fps original video to 60/120fps, preserving motion detail while reducing stutter and ghosting.Frame interpolation focuses on "filling frames" along the time axis. Traditional methods rely on optical flow estimation, first predicting the motion of each pixel between two frames, then interpolating at an intermediate position according to certain rules to generate a new frame. However, in areas of fast motion, multi-object occlusion, or complex texture, optical flow is often not accurate enough, easily producing ghosting, double images, or local deformation. Deep frame interpolation models such as DAIN, RIFE, and FILM simultaneously learn fusion strategies for optical flow, depth, or intermediate features through end-to-end networks, directly outputting interpolated frames, with noticeably improved stability and visual quality in complex scenes. For sports events, action game recordings, and slow-motion creation, frame interpolation can smoothly boost 24/30fps original video to 60/120fps, preserving motion detail while reducing stutter and ghosting.
In engineering practice, super-resolution and frame interpolation are often used together: for low-resolution, low-frame-rate existing content, temporal frame interpolation is performed first, followed by spatial super-resolution, or both are implemented in a unified spatiotemporal network. In terms of deployment, cloud offline processing is suitable for film restoration and platform-level "quality upgrade" services that demand the highest image quality, while device-side real-time inference is more commonly seen in TV boxes, player apps, and gaming/action cameras, requiring low latency through model compression and hardware acceleration. Regardless of the form they take, super-resolution and frame interpolation have become essential infrastructure for the "HD/UHD experience," giving old content a "second life" on new terminals.In engineering practice, super-resolution and frame interpolation are often used together: for low-resolution, low-frame-rate existing content, temporal frame interpolation is performed first, followed by spatial super-resolution, or both are implemented in a unified spatiotemporal network. In terms of deployment, cloud offline processing is suitable for film restoration and platform-level "quality upgrade" services that demand the highest image quality, while device-side real-time inference is more commonly seen in TV boxes, player apps, and gaming/action cameras, requiring low latency through model compression and hardware acceleration. Regardless of the form they take, super-resolution and frame interpolation have become essential infrastructure for the "HD/UHD experience," giving old content a "second life" on new terminals.
If traditional video processing stays more at the level of "image quality and stability," then video understanding and structural analysis begins to answer semantic questions of the type "what is happening in the video": who is doing what, where, for how long, and whether there is abnormal behavior. The goal here is to structurally decompose the video along the time axis: recognize actions and behaviors, detect and track targets, segment foreground and background, partition scenes and shots, and extract high-level semantic signals usable for downstream decision-making, retrieval, and alerting.If traditional video processing stays more at the level of "image quality and stability," then video understanding and structural analysis begins to answer semantic questions of the type "what is happening in the video": who is doing what, where, for how long, and whether there is abnormal behavior. The goal here is to structurally decompose the video along the time axis: recognize actions and behaviors, detect and track targets, segment foreground and background, partition scenes and shots, and extract high-level semantic signals usable for downstream decision-making, retrieval, and alerting.
From a product perspective, this layer of capabilities has already penetrated various intelligent security platforms, sports training analysis systems, smart dashcams, and industrial quality inspection video analysis systems: identifying fighting, falling, loitering, and other anomalies in surveillance; analyzing movement standardization and technical details in sports and fitness scenarios; tracking vehicle and person trajectories and monitoring whether production processes are normal in traffic and industrial environments. Below, we continue to organize these capabilities from three angles — scenarios, principles, and models — and focus on several representative directions in subsequent subsections.From a product perspective, this layer of capabilities has already penetrated various intelligent security platforms, sports training analysis systems, smart dashcams, and industrial quality inspection video analysis systems: identifying fighting, falling, loitering, and other anomalies in surveillance; analyzing movement standardization and technical details in sports and fitness scenarios; tracking vehicle and person trajectories and monitoring whether production processes are normal in traffic and industrial environments. Below, we continue to organize these capabilities from three angles — scenarios, principles, and models — and focus on several representative directions in subsequent subsections.
The key to video understanding and structural analysis is joint modeling of spatial targets and semantics in the temporal dimension:The key to video understanding and structural analysis is joint modeling of spatial targets and semantics in the temporal dimension:
In terms of model selection, video understanding and structural analysis typically adopt a combined architecture of "spatial features + temporal modeling":In terms of model selection, video understanding and structural analysis typically adopt a combined architecture of "spatial features + temporal modeling":
Overall, this layer of capabilities further abstracts video from a "high-quality pixel stream" into a "behavior and event stream," laying the structural foundation for upstream multimodal understanding, retrieval, and decision-making. Below, we expand on three directions: action recognition and behavior analysis, object detection and tracking, and event and anomaly detection.Overall, this layer of capabilities further abstracts video from a "high-quality pixel stream" into a "behavior and event stream," laying the structural foundation for upstream multimodal understanding, retrieval, and decision-making. Below, we expand on three directions: action recognition and behavior analysis, object detection and tracking, and event and anomaly detection.
Action recognition and behavior analysis is concerned with "what the subject is doing within a time window." In security scenarios, this means recognizing behaviors such as "walking, running, falling, fighting" from video; in sports and fitness, it corresponds to more fine-grained actions such as "whether the basketball shot, tennis serve, or squat is standard" and "whether the yoga pose is correct." Technically, early methods mainly relied on 2D convolution + optical flow or handcrafted features, stacking several frames and classifying them as a whole; modern methods more commonly use 3D convolution (I3D, a series of 3D ResNet variants), multi-temporal-scale structures such as SlowFast, or spatiotemporal-attention-based models such as TimeSformer and Video Swin Transformer to jointly model spatial textures and temporal changes.Action recognition and behavior analysis is concerned with "what the subject is doing within a time window." In security scenarios, this means recognizing behaviors such as "walking, running, falling, fighting" from video; in sports and fitness, it corresponds to more fine-grained actions such as "whether the basketball shot, tennis serve, or squat is standard" and "whether the yoga pose is correct." Technically, early methods mainly relied on 2D convolution + optical flow or handcrafted features, stacking several frames and classifying them as a whole; modern methods more commonly use 3D convolution (I3D, a series of 3D ResNet variants), multi-temporal-scale structures such as SlowFast, or spatiotemporal-attention-based models such as TimeSformer and Video Swin Transformer to jointly model spatial textures and temporal changes.
In many scenarios requiring high-precision pose analysis, directly classifying RGB clips is insufficient; human pose estimation and skeletal sequence modeling are also incorporated: 2D/3D keypoints are first extracted from each frame, and the keypoint sequence is then fed into RNNs, temporal convolutions, or GCN/Transformer networks to analyze the temporal structure and spatial coordination of the action. This "pose prior + temporal modeling" approach is more robust to changes in background, lighting, and clothing, and is suitable for applications with high demands on action detail, such as yoga, fitness, and industrial operation compliance assessment.In many scenarios requiring high-precision pose analysis, directly classifying RGB clips is insufficient; human pose estimation and skeletal sequence modeling are also incorporated: 2D/3D keypoints are first extracted from each frame, and the keypoint sequence is then fed into RNNs, temporal convolutions, or GCN/Transformer networks to analyze the temporal structure and spatial coordination of the action. This "pose prior + temporal modeling" approach is more robust to changes in background, lighting, and clothing, and is suitable for applications with high demands on action detail, such as yoga, fitness, and industrial operation compliance assessment.
Single-frame object detection can tell us "what targets are in this frame and where," but many real-world tasks require knowing "where this car / person came from, where they went, and what they did in between." The object detection and tracking module exists precisely to string frame-level detections into continuous temporal trajectories: on one hand, a detector runs on each frame, producing candidate bounding boxes; on the other hand, based on cues such as appearance features (ReID embeddings), motion prediction (Kalman filtering), and spatial overlap, boxes on adjacent frames are matched and associated, yielding multi-object tracking (MOT) results.Single-frame object detection can tell us "what targets are in this frame and where," but many real-world tasks require knowing "where this car / person came from, where they went, and what they did in between." The object detection and tracking module exists precisely to string frame-level detections into continuous temporal trajectories: on one hand, a detector runs on each frame, producing candidate bounding boxes; on the other hand, based on cues such as appearance features (ReID embeddings), motion prediction (Kalman filtering), and spatial overlap, boxes on adjacent frames are matched and associated, yielding multi-object tracking (MOT) results.
In engineering practice, a typical pipeline is: "robust pedestrian / vehicle detection + an association algorithm such as DeepSORT," deployed on surveillance cameras or dashcams, outputting the motion trajectory of each ID in real time. In more complex systems, these trajectories are further combined with regional semantics (lanes, zone divisions) and business logic rules to infer higher-level behavior patterns such as driving against traffic, prolonged loitering, and frequent entry/exit, providing continuous temporal signals for upstream security, traffic flow analysis, and industrial process monitoring.In engineering practice, a typical pipeline is: "robust pedestrian / vehicle detection + an association algorithm such as DeepSORT," deployed on surveillance cameras or dashcams, outputting the motion trajectory of each ID in real time. In more complex systems, these trajectories are further combined with regional semantics (lanes, zone divisions) and business logic rules to infer higher-level behavior patterns such as driving against traffic, prolonged loitering, and frequent entry/exit, providing continuous temporal signals for upstream security, traffic flow analysis, and industrial process monitoring.
In most business scenarios, what truly requires focused attention is often the "minority of anomalies" and "key events": for example, fighting, falling, and gathering in security; abnormal shutdowns or non-compliant operations in industrial production; dangerous driving behavior in traffic. Such events are relatively rare, with high annotation costs and extremely imbalanced samples, posing additional challenges for model construction.In most business scenarios, what truly requires focused attention is often the "minority of anomalies" and "key events": for example, fighting, falling, and gathering in security; abnormal shutdowns or non-compliant operations in industrial production; dangerous driving behavior in traffic. Such events are relatively rare, with high annotation costs and extremely imbalanced samples, posing additional challenges for model construction.
A common approach is to build a temporal anomaly detection module on top of basic action recognition, object tracking, and scene segmentation: either directly learning from a small number of labeled anomaly samples through supervised methods; or using unsupervised/weakly supervised methods to model the motion and behavior distribution of "normal patterns," issuing an alert whenever a new observation significantly deviates from the historical distribution. At the model level, temporal autoencoders, contrastive learning, graph neural networks, or temporal Transformers are combined to uniformly encode spatial relationships and temporal dependencies, thereby capturing more complex group behavior patterns and long-range dependencies.A common approach is to build a temporal anomaly detection module on top of basic action recognition, object tracking, and scene segmentation: either directly learning from a small number of labeled anomaly samples through supervised methods; or using unsupervised/weakly supervised methods to model the motion and behavior distribution of "normal patterns," issuing an alert whenever a new observation significantly deviates from the historical distribution. At the model level, temporal autoencoders, contrastive learning, graph neural networks, or temporal Transformers are combined to uniformly encode spatial relationships and temporal dependencies, thereby capturing more complex group behavior patterns and long-range dependencies.
If video understanding solves the problem of "understanding the video itself clearly," then video + language multimodal tasks focus on "how to use natural language to describe, answer questions about, and retrieve video content," as well as "how to quickly locate key information along the long video timeline based on textual needs." Such tasks require simultaneously processing visual, speech, and text signals: on one hand, extracting visual and audio features from the video; on the other, connecting to the reasoning and generation capabilities of language models, compressing spatiotemporal content into text summaries, Q&A results, and semantic indices suitable for human consumption and machine invocation.If video understanding solves the problem of "understanding the video itself clearly," then video + language multimodal tasks focus on "how to use natural language to describe, answer questions about, and retrieve video content," as well as "how to quickly locate key information along the long video timeline based on textual needs." Such tasks require simultaneously processing visual, speech, and text signals: on one hand, extracting visual and audio features from the video; on the other, connecting to the reasoning and generation capabilities of language models, compressing spatiotemporal content into text summaries, Q&A results, and semantic indices suitable for human consumption and machine invocation.
From a product perspective, this layer of capabilities has already penetrated scenarios such as automatic subtitle and timeline generation for long videos, "smart marking / key clip extraction" on short-video editing platforms, and Q&A assistants for corporate training and meeting videos: users no longer need to "watch from beginning to end," but can directly search, ask questions about, and reorganize video content through natural language. Below, we continue to expand from three angles — scenarios, principles, and models.From a product perspective, this layer of capabilities has already penetrated scenarios such as automatic subtitle and timeline generation for long videos, "smart marking / key clip extraction" on short-video editing platforms, and Q&A assistants for corporate training and meeting videos: users no longer need to "watch from beginning to end," but can directly search, ask questions about, and reorganize video content through natural language. Below, we continue to expand from three angles — scenarios, principles, and models.
The core of video–language multimodal systems is to align temporal visual features with text representations in a unified embedding space, and to perform retrieval, generation, and reasoning on this foundation:The core of video–language multimodal systems is to align temporal visual features with text representations in a unified embedding space, and to perform retrieval, generation, and reasoning on this foundation:
In terms of model form, video–language multimodal tasks have undergone an evolution from "dedicated encoder + simple head" to "unified multimodal large models":In terms of model form, video–language multimodal tasks have undergone an evolution from "dedicated encoder + simple head" to "unified multimodal large models":
Overall, this layer elevates video from "machine understanding" further to the level of "human–machine dialogue and collaboration": users can ask questions about a video as if asking a person, while the system performs complex visual, speech, and language alignment and reasoning behind the scenes.Overall, this layer elevates video from "machine understanding" further to the level of "human–machine dialogue and collaboration": users can ask questions about a video as if asking a person, while the system performs complex visual, speech, and language alignment and reasoning behind the scenes.
For courses, lectures, meetings, and long-form content videos, the most urgent need is often to "quickly know what was said and where the key points are," rather than watching the entire thing from beginning to end. Automatic subtitle and summary systems use a combination of "ASR + text processing + visual assistance" to transcribe audio content into timestamp-aligned text, and then generate structured outlines and concise summaries on that basis, achieving information compression from "hour-long video" to "minute-level reading."For courses, lectures, meetings, and long-form content videos, the most urgent need is often to "quickly know what was said and where the key points are," rather than watching the entire thing from beginning to end. Automatic subtitle and summary systems use a combination of "ASR + text processing + visual assistance" to transcribe audio content into timestamp-aligned text, and then generate structured outlines and concise summaries on that basis, achieving information compression from "hour-long video" to "minute-level reading."
At the implementation level, the ASR module is responsible for stably and reliably producing multilingual transcription and timeline alignment; the text side uses a large language model to correct errors, segment sentences, and semantically reorganize the raw transcription, extracting chapter titles, key information, and question–answer pairs. In some scenarios, visual cues (such as PPT page changes, scene transitions) are also incorporated to assist in demarcating chapter boundaries and key clips, ensuring that the summary structure is more consistent with the rhythm of the actual content.At the implementation level, the ASR module is responsible for stably and reliably producing multilingual transcription and timeline alignment; the text side uses a large language model to correct errors, segment sentences, and semantically reorganize the raw transcription, extracting chapter titles, key information, and question–answer pairs. In some scenarios, visual cues (such as PPT page changes, scene transitions) are also incorporated to assist in demarcating chapter boundaries and key clips, ensuring that the summary structure is more consistent with the rhythm of the actual content.
Building on subtitles and summaries, a further need is the ability to perform Q&A and retrieval on specific video content: for example, "where did this person put the phone in the end," "which part talks about the pricing strategy," "at what minute is this step demonstrated." Such tasks require semantic localization of the query along the time axis: understanding the people, objects, and actions involved in the query, and finding the corresponding clip in the video's temporal representation.Building on subtitles and summaries, a further need is the ability to perform Q&A and retrieval on specific video content: for example, "where did this person put the phone in the end," "which part talks about the pricing strategy," "at what minute is this step demonstrated." Such tasks require semantic localization of the query along the time axis: understanding the people, objects, and actions involved in the query, and finding the corresponding clip in the video's temporal representation.
In concrete terms, a multi-granularity index is typically built offline for the video: multimodal representations (visual + text/speech) are extracted for fixed-length clips, and a vector index or graph structure is established. During online interaction, the user's question is encoded as a text vector and matched against the clip representations in the index to find the most relevant time intervals; then, the content of these clips (keyframe screenshot descriptions, transcription text, etc.) is fed into an LLM together with the question, and the model generates a natural language answer or returns the corresponding time points. For large-scale video libraries, "cross-video retrieval" can be supported under the same mechanism, for example, searching across collections in corporate training knowledge bases or e-commerce product videos.In concrete terms, a multi-granularity index is typically built offline for the video: multimodal representations (visual + text/speech) are extracted for fixed-length clips, and a vector index or graph structure is established. During online interaction, the user's question is encoded as a text vector and matched against the clip representations in the index to find the most relevant time intervals; then, the content of these clips (keyframe screenshot descriptions, transcription text, etc.) is fed into an LLM together with the question, and the model generates a natural language answer or returns the corresponding time points. For large-scale video libraries, "cross-video retrieval" can be supported under the same mechanism, for example, searching across collections in corporate training knowledge bases or e-commerce product videos.
Once the system can stably understand the content and semantic structure of a video, the natural next step is to reversely use these understanding results to assist in creation and editing. Video–language multimodal models can automatically select clips that match the semantics from existing material based on a script or prompt provided by the creator, generating a rough-cut timeline; they can also automatically generate titles, cover copy, and chapter labels based on the video content, and even suggest shot rhythm and soundtrack choices.Once the system can stably understand the content and semantic structure of a video, the natural next step is to reversely use these understanding results to assist in creation and editing. Video–language multimodal models can automatically select clips that match the semantics from existing material based on a script or prompt provided by the creator, generating a rough-cut timeline; they can also automatically generate titles, cover copy, and chapter labels based on the video content, and even suggest shot rhythm and soundtrack choices.
In workflows, such capabilities typically appear in the form of "intelligent recommendations" and "automatic rough cuts": after the creator uploads material, the system automatically completes analysis, storyboarding, and marking, and provides several candidate versions (such as editing plans with different rhythms and durations); the creator can fine-tune on this basis without needing to screen frame by frame from scratch. For enterprise applications, the system can also combine knowledge bases and brand guidelines to ensure that the generated copy, subtitles, and editing style comply with established business requirements and compliance standards.In workflows, such capabilities typically appear in the form of "intelligent recommendations" and "automatic rough cuts": after the creator uploads material, the system automatically completes analysis, storyboarding, and marking, and provides several candidate versions (such as editing plans with different rhythms and durations); the creator can fine-tune on this basis without needing to screen frame by frame from scratch. For enterprise applications, the system can also combine knowledge bases and brand guidelines to ensure that the generated copy, subtitles, and editing style comply with established business requirements and compliance standards.
After possessing stable understanding and structural analysis capabilities, video generation and editing steps into the phase of "actively creating content": no longer just improving image quality or performing structural analysis, but generating entirely new shots based on text scripts, reference images, or existing videos, or performing structural editing and reorganization of original videos. This includes both text-to-video generation from scratch, as well as style transfer, extension, and rearrangement based on existing images/videos, and fine-grained object-level editing and replacement.After possessing stable understanding and structural analysis capabilities, video generation and editing steps into the phase of "actively creating content": no longer just improving image quality or performing structural analysis, but generating entirely new shots based on text scripts, reference images, or existing videos, or performing structural editing and reorganization of original videos. This includes both text-to-video generation from scratch, as well as style transfer, extension, and rearrangement based on existing images/videos, and fine-grained object-level editing and replacement.
In terms of products, this layer of capabilities has already entered the content creation mainstream through a series of products such as Jimeng Video, MiniMax Video, Sora, Runway Gen‑2, Pika, and Kling: advertisements, concept films, animations, and storyboard sequences can be quickly generated without relying on large filming crews and complex post-production; creators can drive shots and styles through natural language scripts; traditional video editing workflows are beginning to deeply integrate with structured generation tools. Below, we continue to organize from the angles of scenarios, principles, and models.In terms of products, this layer of capabilities has already entered the content creation mainstream through a series of products such as Jimeng Video, MiniMax Video, Sora, Runway Gen‑2, Pika, and Kling: advertisements, concept films, animations, and storyboard sequences can be quickly generated without relying on large filming crews and complex post-production; creators can drive shots and styles through natural language scripts; traditional video editing workflows are beginning to deeply integrate with structured generation tools. Below, we continue to organize from the angles of scenarios, principles, and models.
Current mainstream video generation and editing methods mostly use diffusion models or their variants as the core, gradually "denoising" to generate video in a high-dimensional spatiotemporal latent space:Current mainstream video generation and editing methods mostly use diffusion models or their variants as the core, gradually "denoising" to generate video in a high-dimensional spatiotemporal latent space:
Representative models and directions include:Representative models and directions include:
These capabilities do not exist in isolation but are gradually permeating editing and post-production pipelines: from copy to storyboard, storyboard to rough cut, rough cut to stylization and local editing — more and more stages are being driven by "text + structured control."These capabilities do not exist in isolation but are gradually permeating editing and post-production pipelines: from copy to storyboard, storyboard to rough cut, rough cut to stylization and local editing — more and more stages are being driven by "text + structured control."
What text-to-video (Text‑to‑Video) aims to achieve is: the user describes a scene, shot, or story fragment in natural language, and the system automatically generates a coherent video. Compared to image generation, text-to-video adds the challenge of the temporal dimension: not only must image quality and style consistency be maintained at the single-frame level, but the coherence of subject identity, lighting, background, and motion trajectories across frames must also be ensured.What text-to-video (Text‑to‑Video) aims to achieve is: the user describes a scene, shot, or story fragment in natural language, and the system automatically generates a coherent video. Compared to image generation, text-to-video adds the challenge of the temporal dimension: not only must image quality and style consistency be maintained at the single-frame level, but the coherence of subject identity, lighting, background, and motion trajectories across frames must also be ensured.
Typical diffusion-based text-to-video models are first pretrained on large-scale video–text paired data: a text encoder extracts semantic conditions, and a video decoder repeatedly denoises a "noisy video" in latent space, gradually converging to spatiotemporal signals consistent with the text. In this process, structures such as temporal attention, 3D convolution, or 4D representations explicitly build temporal dependencies into the network to avoid problems such as "inter-frame jumping" and "character resetting." Some systems also support control over shot motion (push, pull, pan, tilt) and composition rhythm, making the generated results closer to real cinematographic language.Typical diffusion-based text-to-video models are first pretrained on large-scale video–text paired data: a text encoder extracts semantic conditions, and a video decoder repeatedly denoises a "noisy video" in latent space, gradually converging to spatiotemporal signals consistent with the text. In this process, structures such as temporal attention, 3D convolution, or 4D representations explicitly build temporal dependencies into the network to avoid problems such as "inter-frame jumping" and "character resetting." Some systems also support control over shot motion (push, pull, pan, tilt) and composition rhythm, making the generated results closer to real cinematographic language.
Another important route is generation and editing based on existing images or videos: for example, "bringing to life" an illustration or concept design, stylizing real-person video into anime, or changing the background, adjusting weather and time while keeping the structure unchanged. Technically, such methods often add a "reference channel" to the diffusion process: the input image or video is encoded as features, participating in denoising as a condition or initial state, while mechanisms such as masks and explicit geometric constraints control "which regions can be changed and which must be preserved."Another important route is generation and editing based on existing images or videos: for example, "bringing to life" an illustration or concept design, stylizing real-person video into anime, or changing the background, adjusting weather and time while keeping the structure unchanged. Technically, such methods often add a "reference channel" to the diffusion process: the input image or video is encoded as features, participating in denoising as a condition or initial state, while mechanisms such as masks and explicit geometric constraints control "which regions can be changed and which must be preserved."
For style transfer scenarios, the model redraws textures and lighting while preserving the original motion and composition, matching the target style; for video extension and reorganization, new frames are "continued" at the temporal ends or in the middle, achieving horizontal/vertical scene expansion, viewpoint orbiting, or plot supplementation. Such capabilities are very suitable for integration with traditional editing workflows: the editor first provides key shots and rhythm, and the model then automatically generates transitions and variations between these "anchor points."For style transfer scenarios, the model redraws textures and lighting while preserving the original motion and composition, matching the target style; for video extension and reorganization, new frames are "continued" at the temporal ends or in the middle, achieving horizontal/vertical scene expansion, viewpoint orbiting, or plot supplementation. Such capabilities are very suitable for integration with traditional editing workflows: the editor first provides key shots and rhythm, and the model then automatically generates transitions and variations between these "anchor points."
In many business scenarios, fully regenerating video is not a hard requirement; what is more critical is performing fine, controllable structured editing on existing footage: for example, face swapping, modifying mouth shapes, erasing unwanted objects, replacing advertising content, or rearranging shot order based on a text script. Structured video editing develops along this line of thinking: building on video understanding, object-level segmentation, tracking, and parametric representations are introduced, so that editing operations can be stably bound to specific targets and time periods.In many business scenarios, fully regenerating video is not a hard requirement; what is more critical is performing fine, controllable structured editing on existing footage: for example, face swapping, modifying mouth shapes, erasing unwanted objects, replacing advertising content, or rearranging shot order based on a text script. Structured video editing develops along this line of thinking: building on video understanding, object-level segmentation, tracking, and parametric representations are introduced, so that editing operations can be stably bound to specific targets and time periods.
Face swapping and lip-sync are the most typical applications in this direction: the model needs to map the target person's identity onto the original video's performance while ensuring natural and coherent head pose and overall expression, and precisely control mouth shape movements based on the new speech signal. Object erasure / replacement relies on high-quality segmentation and spatiotemporal inpainting: first segment and remove the target object in each frame, then fill the hole using neighboring frames and contextual texture, avoiding obvious "patching" traces. Text-driven editing aligns the "script structure" with the video timeline, automatically selecting and stitching clips that match the script semantics, achieving higher-level automated editing.Face swapping and lip-sync are the most typical applications in this direction: the model needs to map the target person's identity onto the original video's performance while ensuring natural and coherent head pose and overall expression, and precisely control mouth shape movements based on the new speech signal. Object erasure / replacement relies on high-quality segmentation and spatiotemporal inpainting: first segment and remove the target object in each frame, then fill the hole using neighboring frames and contextual texture, avoiding obvious "patching" traces. Text-driven editing aligns the "script structure" with the video timeline, automatically selecting and stitching clips that match the script semantics, achieving higher-level automated editing.
Digital Human / Avatar can be seen as a "system-level integration" of video generation, speech synthesis, multimodal understanding, and graphics rendering: it is not just about generating a segment of video, but about continuously and controllably driving a virtual figure to "speak, make expressions, and gesture" based on text or speech input, and in more and more scenarios, achieving near-real-time or even real-time interaction. Compared to general video generation, digital humans emphasize three points more: long-term consistency of identity and appearance, fine alignment of speech–expression–motion, and the real-time performance and stability of the end-to-end system.Digital Human / Avatar can be seen as a "system-level integration" of video generation, speech synthesis, multimodal understanding, and graphics rendering: it is not just about generating a segment of video, but about continuously and controllably driving a virtual figure to "speak, make expressions, and gesture" based on text or speech input, and in more and more scenarios, achieving near-real-time or even real-time interaction. Compared to general video generation, digital humans emphasize three points more: long-term consistency of identity and appearance, fine alignment of speech–expression–motion, and the real-time performance and stability of the end-to-end system.
From a product perspective, digital humans have already widely appeared in scenarios such as content production platforms, virtual customer service / smart reception / virtual guided tours, education and training and online classrooms, brand virtual IP / virtual idols, and virtual streamer / digital avatar tools for creators: enterprises can batch-produce video content with fixed appearances and styles, government and enterprise services can use virtual receptionists to serve users 24/7, and individual creators can consistently produce "person-on-camera" videos without ever showing their face. Below, we continue to organize from three dimensions — scenarios, principles, and models — and expand on three directions in subsequent subsections: driving and expression, appearance and video generation, and real-time interaction and system integration.From a product perspective, digital humans have already widely appeared in scenarios such as content production platforms, virtual customer service / smart reception / virtual guided tours, education and training and online classrooms, brand virtual IP / virtual idols, and virtual streamer / digital avatar tools for creators: enterprises can batch-produce video content with fixed appearances and styles, government and enterprise services can use virtual receptionists to serve users 24/7, and individual creators can consistently produce "person-on-camera" videos without ever showing their face. Below, we continue to organize from three dimensions — scenarios, principles, and models — and expand on three directions in subsequent subsections: driving and expression, appearance and video generation, and real-time interaction and system integration.
A digital human system is essentially a multimodal pipeline of "speech / text driving + appearance modeling + video / rendering output," with slight differences between offline and real-time scenarios but similar core components:A digital human system is essentially a multimodal pipeline of "speech / text driving + appearance modeling + video / rendering output," with slight differences between offline and real-time scenarios but similar core components:
In terms of specific models, digital human systems often comprehensively use multiple types of specialized models and general multimodal models:In terms of specific models, digital human systems often comprehensively use multiple types of specialized models and general multimodal models:
Taken together, digital humans are both a set of models and a complete system: they integrate language understanding, speech, visual generation, and real-time inference to present an interactive virtual character "on screen." Below, we expand on three directions: driving and expression, appearance and video generation, and real-time interaction and system integration.Taken together, digital humans are both a set of models and a complete system: they integrate language understanding, speech, visual generation, and real-time inference to present an interactive virtual character "on screen." Below, we expand on three directions: driving and expression, appearance and video generation, and real-time interaction and system integration.
In the digital human pipeline, driving and expression is responsible for answering a core question: given a script or speech, what mouth shape, expression, and head-and-shoulder movements should the virtual figure present in each frame. This includes both offline batch production scenarios and responses to real-time dialogue.In the digital human pipeline, driving and expression is responsible for answering a core question: given a script or speech, what mouth shape, expression, and head-and-shoulder movements should the virtual figure present in each frame. This includes both offline batch production scenarios and responses to real-time dialogue.
In offline content production, a common pipeline is "text script → TTS → speech-driven": the business side provides the narration script, the TTS module generates speech in the target timbre (such as a brand's virtual spokesperson), and the speech features are then fed into a "speech → motion" model. Wav2Lip-type models are an important representative of this stage:In offline content production, a common pipeline is "text script → TTS → speech-driven": the business side provides the narration script, the TTS module generates speech in the target timbre (such as a brand's virtual spokesperson), and the speech features are then fed into a "speech → motion" model. Wav2Lip-type models are an important representative of this stage:
Compared to earlier pure lip-sync approaches, the new generation of speech-driven models (such as MuseTalk-type methods) further extend to full-face expressions and head pose:Compared to earlier pure lip-sync approaches, the new generation of speech-driven models (such as MuseTalk-type methods) further extend to full-face expressions and head pose:
At a higher dimension, driving and expression can also incorporate external control signals: for example, using pose skeletons, gesture trajectories, and gaze direction as additional inputs, enabling the digital human to imitate the style of a specific speaker, or to execute predefined action templates based on "directive actions" in the script (such as "point to the screen," "open both hands"). Whether it is local mouth-shape driving like Wav2Lip, or more full-body expression modeling like MuseTalk / real-time skeleton driving, they together achieve continuous mapping from speech/text to facial and upper-body movements, and are the key link that makes a digital human "look like it's seriously speaking."At a higher dimension, driving and expression can also incorporate external control signals: for example, using pose skeletons, gesture trajectories, and gaze direction as additional inputs, enabling the digital human to imitate the style of a specific speaker, or to execute predefined action templates based on "directive actions" in the script (such as "point to the screen," "open both hands"). Whether it is local mouth-shape driving like Wav2Lip, or more full-body expression modeling like MuseTalk / real-time skeleton driving, they together achieve continuous mapping from speech/text to facial and upper-body movements, and are the key link that makes a digital human "look like it's seriously speaking."
The driving pipeline solves "how to move," while appearance and video generation determines "who is moving, where they are moving, and in what style." This includes both high-fidelity photorealistic digital humans and stylized figures such as 2D anime, cartoon, and low-poly avatars, as well as different technical choices for real-time and offline rendering.The driving pipeline solves "how to move," while appearance and video generation determines "who is moving, where they are moving, and in what style." This includes both high-fidelity photorealistic digital humans and stylized figures such as 2D anime, cartoon, and low-poly avatars, as well as different technical choices for real-time and offline rendering.
In 2D portrait and illustration scenarios, the typical approach is to train a Talking Head generation model based on a small number of reference images and short videos:In 2D portrait and illustration scenarios, the typical approach is to train a Talking Head generation model based on a small number of reference images and short videos:
In scenarios pursuing higher realism, freer viewpoints, and multi-camera switching, an increasing number of solutions adopt digital human modeling based on NeRF / 4D representations (such as ER‑NeRF-type methods):In scenarios pursuing higher realism, freer viewpoints, and multi-camera switching, an increasing number of solutions adopt digital human modeling based on NeRF / 4D representations (such as ER‑NeRF-type methods):
In businesses emphasizing cross-platform deployment and real-time performance, lightweight solutions such as Ultralight‑Digital‑Human are also adopted:In businesses emphasizing cross-platform deployment and real-time performance, lightweight solutions such as Ultralight‑Digital‑Human are also adopted:
At the level of complete video production, appearance and video generation must also be combined with backgrounds, props, and cinematic language. A common workflow is:At the level of complete video production, appearance and video generation must also be combined with backgrounds, props, and cinematic language. A common workflow is:
This makes the digital human not just a "talking head," but a "character" that can naturally blend into various program and content formats.This makes the digital human not just a "talking head," but a "character" that can naturally blend into various program and content formats.
As ASR, TTS, LLM, and lightweight video generation models mature, more and more digital human systems are moving from offline batch video production to real-time interaction: the user speaks or types at the terminal, and within a few hundred milliseconds to a few seconds, the digital human on screen "understands — thinks — responds — speaks," creating an experience similar to a real customer service agent / guide / host. The key here is not just the models themselves, but also how to compress the multimodal pipeline to acceptable end-to-end latency.As ASR, TTS, LLM, and lightweight video generation models mature, more and more digital human systems are moving from offline batch video production to real-time interaction: the user speaks or types at the terminal, and within a few hundred milliseconds to a few seconds, the digital human on screen "understands — thinks — responds — speaks," creating an experience similar to a real customer service agent / guide / host. The key here is not just the models themselves, but also how to compress the multimodal pipeline to acceptable end-to-end latency.
In a typical real-time digital human closed loop:In a typical real-time digital human closed loop:
To provide a consistent experience across multiple terminals, the system also needs to make careful trade-offs among latency, bandwidth, and compute:To provide a consistent experience across multiple terminals, the system also needs to make careful trade-offs among latency, bandwidth, and compute:
On the model side, real-time digital humans also impose additional requirements on architectural design:On the model side, real-time digital humans also impose additional requirements on architectural design:
At the system integration level, real-time digital humans often also need to be tightly bound to business knowledge, persona settings, and dialogue strategies:At the system integration level, real-time digital humans often also need to be tightly bound to business knowledge, persona settings, and dialogue strategies:
Overall, with the addition of models such as Wav2Lip, MuseTalk, ER‑NeRF, and Ultralight‑Digital‑Human specifically designed for lip-sync, expression driving, and real-time rendering, digital humans are accelerating their evolution from "offline video template tools" into virtual entities that can respond in real time, have stable personalities and professional knowledge, becoming the most comprehensive and application-rich component of the video technology system.Overall, with the addition of models such as Wav2Lip, MuseTalk, ER‑NeRF, and Ultralight‑Digital‑Human specifically designed for lip-sync, expression driving, and real-time rendering, digital humans are accelerating their evolution from "offline video template tools" into virtual entities that can respond in real time, have stable personalities and professional knowledge, becoming the most comprehensive and application-rich component of the video technology system.
In the previous sections on vision and structured modeling, we mostly thought about problems in "static" spaces: an image, a record, a piece of text. In real business, however, a large portion of core metrics evolve over time: sales and traffic fluctuate daily, server load and sensor readings change every second, and financial prices and macro indicators continually adjust under the influence of policies and events. The Time Series & Sequential Decision layer focuses on: forecasting the future along the time axis, identifying anomalies, characterizing structural breaks, and on that basis making forward-looking decisions and control actions.In the previous sections on vision and structured modeling, we mostly thought about problems in "static" spaces: an image, a record, a piece of text. In real business, however, a large portion of core metrics evolve over time: sales and traffic fluctuate daily, server load and sensor readings change every second, and financial prices and macro indicators continually adjust under the influence of policies and events. The Time Series & Sequential Decision layer focuses on: forecasting the future along the time axis, identifying anomalies, characterizing structural breaks, and on that basis making forward-looking decisions and control actions.
From a product perspective, such capabilities span critical functions like operations, planning, risk control, and scheduling: metric forecasting modules embedded in traditional BI/reporting systems, demand forecasting and safety stock recommendations in financial and supply chain planning tools, macro correlation analysis and causal mining in quantitative research software, traffic and capacity forecasting on e-commerce and mobility platforms, and metric anomaly detection and alerting in AIOps for operations. These are all typical product manifestations of this layer. Below, we elaborate from four directions: classical statistical methods, deep learning time series modeling, anomaly & change point detection, and spatio-temporal sequence modeling.From a product perspective, such capabilities span critical functions like operations, planning, risk control, and scheduling: metric forecasting modules embedded in traditional BI/reporting systems, demand forecasting and safety stock recommendations in financial and supply chain planning tools, macro correlation analysis and causal mining in quantitative research software, traffic and capacity forecasting on e-commerce and mobility platforms, and metric anomaly detection and alerting in AIOps for operations. These are all typical product manifestations of this layer. Below, we elaborate from four directions: classical statistical methods, deep learning time series modeling, anomaly & change point detection, and spatio-temporal sequence modeling.
In many business contexts, "time" is the natural backbone: sales vary by day/week, website traffic fluctuates with campaigns, equipment load follows user behavior, and sensor readings reflect subtle changes in system state. Classical statistical time series modeling leverages relatively interpretable and analyzable statistical models on such temporal structures to answer three core questions: What will happen in the future? How are variables related to each other? What is the current state of the system? Although deep learning has emerged prominently in many scenarios, traditional methods like ARIMA, cointegration analysis, and Kalman filtering still serve long-term roles in finance, supply chain, operations, and risk control, and often act as the "baseline" and interpretation tool for more complex systems.In many business contexts, "time" is the natural backbone: sales vary by day/week, website traffic fluctuates with campaigns, equipment load follows user behavior, and sensor readings reflect subtle changes in system state. Classical statistical time series modeling leverages relatively interpretable and analyzable statistical models on such temporal structures to answer three core questions: What will happen in the future? How are variables related to each other? What is the current state of the system? Although deep learning has emerged prominently in many scenarios, traditional methods like ARIMA, cointegration analysis, and Kalman filtering still serve long-term roles in finance, supply chain, operations, and risk control, and often act as the "baseline" and interpretation tool for more complex systems.
From an application perspective, classical time series models are widely present in metric forecasting modules of traditional BI/reporting systems, financial and supply chain planning tools, and various quantitative research software. They can provide future prediction intervals for single or multiple time series, analyze co-movement and long-run equilibrium relationships among macro indicators, and estimate trajectories and hidden states through state-space modeling. Below, we organize the typical usage of these methods along three dimensions — scenarios, principles, and models — and then elaborate on each specific direction.From an application perspective, classical time series models are widely present in metric forecasting modules of traditional BI/reporting systems, financial and supply chain planning tools, and various quantitative research software. They can provide future prediction intervals for single or multiple time series, analyze co-movement and long-run equilibrium relationships among macro indicators, and estimate trajectories and hidden states through state-space modeling. Below, we organize the typical usage of these methods along three dimensions — scenarios, principles, and models — and then elaborate on each specific direction.
Classical time series methods are generally based on the idea of statistical assumptions + parametric structure:Classical time series methods are generally based on the idea of statistical assumptions + parametric structure:
The model family for such methods is relatively well-defined and structurally clear, facilitating interpretation and tuning:The model family for such methods is relatively well-defined and structurally clear, facilitating interpretation and tuning:
Taken together, the strengths of classical time series modeling lie in interpretability, diagnosability, and engineering controllability: the modeling workflow, hypothesis testing, and residual analysis all have mature standards, making them easy to integrate into existing BI and planning systems. Below, we elaborate on three directions: univariate/multivariate forecasting, cointegration and causality, and state-space modeling.Taken together, the strengths of classical time series modeling lie in interpretability, diagnosability, and engineering controllability: the modeling workflow, hypothesis testing, and residual analysis all have mature standards, making them easy to integrate into existing BI and planning systems. Below, we elaborate on three directions: univariate/multivariate forecasting, cointegration and causality, and state-space modeling.
In the most typical business scenarios, the first thing we face is one or several metric curves ordered by time: for example, daily sales of a product, hourly PV of a website, per-minute CPU usage of a server room, or per-second readings from a device sensor. The goal is to provide short- to medium-term forecasts based on historical patterns and to give reasonable confidence intervals. The AR/MA/ARMA/ARIMA/SARIMA family of models is the standard toolkit designed for this purpose.In the most typical business scenarios, the first thing we face is one or several metric curves ordered by time: for example, daily sales of a product, hourly PV of a website, per-minute CPU usage of a server room, or per-second readings from a device sensor. The goal is to provide short- to medium-term forecasts based on historical patterns and to give reasonable confidence intervals. The AR/MA/ARMA/ARIMA/SARIMA family of models is the standard toolkit designed for this purpose.
For univariate series, ARIMA-type models assume that "the current value is linearly determined by past values and random disturbances over a number of lags," and eliminate trends and seasonality by differencing and seasonal differencing to make the series stationary:For univariate series, ARIMA-type models assume that "the current value is linearly determined by past values and random disturbances over a number of lags," and eliminate trends and seasonality by differencing and seasonal differencing to make the series stationary:
In engineering practice, one typically first performs stationarity tests (e.g., ADF), examines ACF/PACF plots, and then selects reasonable orders via information criteria (AIC/BIC) and residual diagnostics. For metrics with pronounced seasonality (e.g., e-commerce daily sales, holiday traffic), SARIMA modeling is especially suitable, and incorporating holiday features or exogenous variables can further improve forecasting performance.In engineering practice, one typically first performs stationarity tests (e.g., ADF), examines ACF/PACF plots, and then selects reasonable orders via information criteria (AIC/BIC) and residual diagnostics. For metrics with pronounced seasonality (e.g., e-commerce daily sales, holiday traffic), SARIMA modeling is especially suitable, and incorporating holiday features or exogenous variables can further improve forecasting performance.
When we wish to model multiple related time series jointly, we can introduce multivariate time series models. The representative method is VAR (Vector AutoRegression) and its variants. VAR treats multiple series as a joint vector, using their own and each other's lagged terms to jointly explain current values, thereby capturing mutual influences among different indicators. For example, in macroeconomic analysis, one can include GDP growth, inflation rate, interest rates, and exchange rates in a single VAR model to study impulse responses and transmission pathways; in business operations, VAR can also describe "how traffic changes in one channel affect other channels" or "the dynamic relationship between promotion intensity and sales," providing references for resource allocation.When we wish to model multiple related time series jointly, we can introduce multivariate time series models. The representative method is VAR (Vector AutoRegression) and its variants. VAR treats multiple series as a joint vector, using their own and each other's lagged terms to jointly explain current values, thereby capturing mutual influences among different indicators. For example, in macroeconomic analysis, one can include GDP growth, inflation rate, interest rates, and exchange rates in a single VAR model to study impulse responses and transmission pathways; in business operations, VAR can also describe "how traffic changes in one channel affect other channels" or "the dynamic relationship between promotion intensity and sales," providing references for resource allocation.
In terms of product form, this type of univariate/multivariate forecasting capability is typically embedded in forecasting functions of traditional BI/reporting systems, and financial and supply chain planning tools: the user selects one or several time series, and the system automatically completes modeling and forecasting, providing prediction intervals, residual analysis, and model diagnostic reports to support decision-making without requiring the user to deeply understand all the mathematical details behind the decisions.In terms of product form, this type of univariate/multivariate forecasting capability is typically embedded in forecasting functions of traditional BI/reporting systems, and financial and supply chain planning tools: the user selects one or several time series, and the system automatically completes modeling and forecasting, providing prediction intervals, residual analysis, and model diagnostic reports to support decision-making without requiring the user to deeply understand all the mathematical details behind the decisions.
In economics and finance, many time series appear to be random walks on the surface, but over longer time scales, there exists some kind of stable long-run equilibrium relationship. Typical examples include exchange rates and interest rate spreads, stock indices and macro earnings, commodity prices and cost indices. Individually, each series may be non-stationary; yet some linear combination oscillates around a stable level in the long run. This phenomenon is called cointegration, and it provides important clues for understanding the structural relationships among macro indicators.In economics and finance, many time series appear to be random walks on the surface, but over longer time scales, there exists some kind of stable long-run equilibrium relationship. Typical examples include exchange rates and interest rate spreads, stock indices and macro earnings, commodity prices and cost indices. Individually, each series may be non-stationary; yet some linear combination oscillates around a stable level in the long run. This phenomenon is called cointegration, and it provides important clues for understanding the structural relationships among macro indicators.
In engineering practice, cointegration analysis typically involves several steps:In engineering practice, cointegration analysis typically involves several steps:
Related to cointegration is the Granger causality test. It is not causality in the strict philosophical sense, but a statistical definition based on predictive power: if the historical information of variable X can significantly improve the prediction accuracy of variable Y, then "X Granger-causes Y." By comparing prediction errors with and without the lagged terms of a certain variable within a VAR or regression framework, one can assess the directional influence among different macro or market indicators. In quantitative research and macro analysis, this type of test is often used to identify potential leading indicators, construct factors, or validate strategy hypotheses.Related to cointegration is the Granger causality test. It is not causality in the strict philosophical sense, but a statistical definition based on predictive power: if the historical information of variable X can significantly improve the prediction accuracy of variable Y, then "X Granger-causes Y." By comparing prediction errors with and without the lagged terms of a certain variable within a VAR or regression framework, one can assess the directional influence among different macro or market indicators. In quantitative research and macro analysis, this type of test is often used to identify potential leading indicators, construct factors, or validate strategy hypotheses.
From a product perspective, cointegration and causality analysis more often appear in quantitative research software, macroeconomic analysis platforms, and financial research tools. They help researchers extract relatively robust structural relationships from piles of time series and map these relationships to higher-level business concepts (such as "the long-run constraint of interest rates on exchange rates" or "spread mean-reversion among different assets"), serving as an important basis for strategy design and risk management.From a product perspective, cointegration and causality analysis more often appear in quantitative research software, macroeconomic analysis platforms, and financial research tools. They help researchers extract relatively robust structural relationships from piles of time series and map these relationships to higher-level business concepts (such as "the long-run constraint of interest rates on exchange rates" or "spread mean-reversion among different assets"), serving as an important basis for strategy design and risk management.
In many real-world systems, the time series we observe is merely a noise-contaminated surface manifestation, while what we are truly interested in is the underlying "system state" evolving over time: for example, the true position and velocity of a vehicle, the health status of equipment, or the latent behavioral patterns of users. In such cases, if we merely do ARIMA-style modeling on the observed series, it is difficult to fully leverage our understanding of the system structure. State-Space Models are proposed precisely for this kind of "hidden state + noisy observation" problem.In many real-world systems, the time series we observe is merely a noise-contaminated surface manifestation, while what we are truly interested in is the underlying "system state" evolving over time: for example, the true position and velocity of a vehicle, the health status of equipment, or the latent behavioral patterns of users. In such cases, if we merely do ARIMA-style modeling on the observed series, it is difficult to fully leverage our understanding of the system structure. State-Space Models are proposed precisely for this kind of "hidden state + noisy observation" problem.
A state-space model typically consists of two parts:A state-space model typically consists of two parts:
Under the linear Gaussian assumption, this framework can achieve recursive estimation and prediction of the state through the Kalman Filter and Smoother: each step is divided into two major phases — "prediction" and "update" — combining the state distribution from the previous moment with the current observation to obtain a new state estimate. This is extremely common in navigation and localization (e.g., trajectory estimation, target tracking), financial time series (e.g., volatility estimation), and equipment state estimation (e.g., health monitoring, remaining useful life prediction).Under the linear Gaussian assumption, this framework can achieve recursive estimation and prediction of the state through the Kalman Filter and Smoother: each step is divided into two major phases — "prediction" and "update" — combining the state distribution from the previous moment with the current observation to obtain a new state estimate. This is extremely common in navigation and localization (e.g., trajectory estimation, target tracking), financial time series (e.g., volatility estimation), and equipment state estimation (e.g., health monitoring, remaining useful life prediction).
Adjacent to continuous state-space models is the Hidden Markov Model (HMM). HMM assumes that the system transitions over time among a number of discrete hidden states, with different probability distributions for generating observed data under each hidden state. Through the forward-backward algorithm and the Viterbi algorithm, HMM can estimate the hidden state sequence, compute the probability of an observation sequence, and predict the next state and observation. HMM was widely used early on in speech recognition and text tagging, and is also commonly used for simple behavioral pattern recognition and event sequence modeling. It still has advantages in certain industrial and financial scenarios — interpretable structure, stable training, and easy integration with domain experience.Adjacent to continuous state-space models is the Hidden Markov Model (HMM). HMM assumes that the system transitions over time among a number of discrete hidden states, with different probability distributions for generating observed data under each hidden state. Through the forward-backward algorithm and the Viterbi algorithm, HMM can estimate the hidden state sequence, compute the probability of an observation sequence, and predict the next state and observation. HMM was widely used early on in speech recognition and text tagging, and is also commonly used for simple behavioral pattern recognition and event sequence modeling. It still has advantages in certain industrial and financial scenarios — interpretable structure, stable training, and easy integration with domain experience.
At the system level, state-space modeling, Kalman filtering, and HMM often serve as the underlying modules for trajectory estimation, equipment state estimation, and financial and engineering control systems, encapsulated within larger toolchains. They may not be directly exposed to end users, but behind products in navigation, target tracking, industrial control, and risk measurement, they have long played the role of "invisible engines."At the system level, state-space modeling, Kalman filtering, and HMM often serve as the underlying modules for trajectory estimation, equipment state estimation, and financial and engineering control systems, encapsulated within larger toolchains. They may not be directly exposed to end users, but behind products in navigation, target tracking, industrial control, and risk measurement, they have long played the role of "invisible engines."
As data scale and scenario complexity continue to rise, classical models that rely solely on linearity and stationarity assumptions increasingly appear "inadequate" in many applications: a large number of nonlinear patterns, long-span dependencies, complex multivariate interactions, sudden behaviors, and period superposition characteristics require more flexible, higher-capacity model structures. Deep learning time series modeling has developed against this backdrop: from RNN/LSTM/GRU, to Temporal CNN/TCN, to time-series-specific Transformers, hybrid and hierarchical models — together they form the main toolkit for modern time series forecasting and modeling.As data scale and scenario complexity continue to rise, classical models that rely solely on linearity and stationarity assumptions increasingly appear "inadequate" in many applications: a large number of nonlinear patterns, long-span dependencies, complex multivariate interactions, sudden behaviors, and period superposition characteristics require more flexible, higher-capacity model structures. Deep learning time series modeling has developed against this backdrop: from RNN/LSTM/GRU, to Temporal CNN/TCN, to time-series-specific Transformers, hybrid and hierarchical models — together they form the main toolkit for modern time series forecasting and modeling.
From an application perspective, deep time series models have been widely deployed in e-commerce traffic & sales forecasting platforms, supply-demand/capacity/scheduling forecasting systems, cloud resource load forecasting and capacity planning tools, used to provide unified and flexible forecasting solutions under complex structures spanning multiple categories, stores, cities, and even business lines. Compared to classical models, they emphasize "end-to-end representation learning" and "global pattern modeling," and are better at handling long sequences, high-dimensional, and multivariate scenarios. Below, we likewise elaborate along three dimensions: scenarios, principles, and models.From an application perspective, deep time series models have been widely deployed in e-commerce traffic & sales forecasting platforms, supply-demand/capacity/scheduling forecasting systems, cloud resource load forecasting and capacity planning tools, used to provide unified and flexible forecasting solutions under complex structures spanning multiple categories, stores, cities, and even business lines. Compared to classical models, they emphasize "end-to-end representation learning" and "global pattern modeling," and are better at handling long sequences, high-dimensional, and multivariate scenarios. Below, we likewise elaborate along three dimensions: scenarios, principles, and models.
The core of deep time series models lies in automatically learning multi-scale patterns and long-term dependencies from historical sequences and covariates:The core of deep time series models lies in automatically learning multi-scale patterns and long-term dependencies from historical sequences and covariates:
In terms of concrete implementation, deep time series modeling has produced a series of representative architectures:In terms of concrete implementation, deep time series modeling has produced a series of representative architectures:
Below, we elaborate on three directions: deep sequence models, convolutional and Transformer models, and hybrid and hierarchical modeling.Below, we elaborate on three directions: deep sequence models, convolutional and Transformer models, and hybrid and hierarchical modeling.
In the early days of deep learning entering the time series domain, RNN/LSTM/GRU were the most natural choice. Similar to text and speech modeling, they "remember" historical information by passing hidden states between time steps, allowing the capture of more complex nonlinearities and long-term dependencies than traditional linear models. For a single or a small number of time series, simple LSTM/GRU can achieve decent forecasting results given sufficient data; in large-scale multi-sequence scenarios, one can adopt shared-parameter RNN/LSTM/GRU models, jointly training on all sequences to learn universal temporal patterns.In the early days of deep learning entering the time series domain, RNN/LSTM/GRU were the most natural choice. Similar to text and speech modeling, they "remember" historical information by passing hidden states between time steps, allowing the capture of more complex nonlinearities and long-term dependencies than traditional linear models. For a single or a small number of time series, simple LSTM/GRU can achieve decent forecasting results given sufficient data; in large-scale multi-sequence scenarios, one can adopt shared-parameter RNN/LSTM/GRU models, jointly training on all sequences to learn universal temporal patterns.
Building on this, autoregressive probabilistic models like DeepAR provide a standard framework for deep time series modeling: they feed historical observations and covariates into a shared RNN/LSTM/GRU network, output the parameters of the conditional distribution of the series value at each time step (e.g., Gaussian, negative binomial, etc.), and achieve end-to-end probabilistic forecasting through maximum likelihood training. This design enables the model to naturally generate prediction intervals, handle irregular scales and multi-series mixing, and facilitates deployment in scenarios such as e-commerce sales and demand forecasting.Building on this, autoregressive probabilistic models like DeepAR provide a standard framework for deep time series modeling: they feed historical observations and covariates into a shared RNN/LSTM/GRU network, output the parameters of the conditional distribution of the series value at each time step (e.g., Gaussian, negative binomial, etc.), and achieve end-to-end probabilistic forecasting through maximum likelihood training. This design enables the model to naturally generate prediction intervals, handle irregular scales and multi-series mixing, and facilitates deployment in scenarios such as e-commerce sales and demand forecasting.
However, RNN-type models have typical problems: gradient decay over long sequences, and inability to fully parallelize during training. Although gating mechanisms (LSTM/GRU) alleviate some of these issues, training and inference efficiency remain factors to weigh for particularly long time spans and high-frequency data. This has also prompted both industry and academia to explore more parallel-friendly structures, such as TCN and Transformer.However, RNN-type models have typical problems: gradient decay over long sequences, and inability to fully parallelize during training. Although gating mechanisms (LSTM/GRU) alleviate some of these issues, training and inference efficiency remain factors to weigh for particularly long time spans and high-frequency data. This has also prompted both industry and academia to explore more parallel-friendly structures, such as TCN and Transformer.
To address the efficiency and stability issues of RNNs on long sequences, Temporal CNN / TCN introduced one-dimensional convolution and dilated convolution to model temporal dependencies: by stacking multiple layers of causal convolution and progressively expanding the receptive field layer by layer, it achieves modeling of long-range history without breaking temporal causality. Compared to RNNs, TCNs can be highly parallel during training and have shorter gradient propagation paths, thus excelling in training stability and efficiency, making them suitable for high-frequency data and industrial time series forecasting scenarios that require large receptive fields.To address the efficiency and stability issues of RNNs on long sequences, Temporal CNN / TCN introduced one-dimensional convolution and dilated convolution to model temporal dependencies: by stacking multiple layers of causal convolution and progressively expanding the receptive field layer by layer, it achieves modeling of long-range history without breaking temporal causality. Compared to RNNs, TCNs can be highly parallel during training and have shorter gradient propagation paths, thus excelling in training stability and efficiency, making them suitable for high-frequency data and industrial time series forecasting scenarios that require large receptive fields.
At a higher level of complexity, Transformers and time-series-specific structures have become the protagonists of long-sequence, multivariate time series modeling in recent years. Directly using standard Transformers encounters the problem of computational complexity growing quadratically with sequence length, so a series of time-series-oriented adaptations have emerged:At a higher level of complexity, Transformers and time-series-specific structures have become the protagonists of long-sequence, multivariate time series modeling in recent years. Directly using standard Transformers encounters the problem of computational complexity growing quadratically with sequence length, so a series of time-series-oriented adaptations have emerged:
These models are often particularly well-suited for complex time series scenarios with long sequences, multiple variables, and high-dimensional covariates, such as large-scale cloud resource loads, multi-region energy demand, and multi-channel traffic forecasting. They can simultaneously model multidimensional inputs, static features, and time-dependent variables within a unified architecture, and provide some clues for subsequent interpretation and diagnosis through attention weights.These models are often particularly well-suited for complex time series scenarios with long sequences, multiple variables, and high-dimensional covariates, such as large-scale cloud resource loads, multi-region energy demand, and multi-channel traffic forecasting. They can simultaneously model multidimensional inputs, static features, and time-dependent variables within a unified architecture, and provide some clues for subsequent interpretation and diagnosis through attention weights.
In real business, time series are rarely "isolated": they often have clear hierarchical structures and shared patterns — for example, store/city/region/national sales hierarchies, SKU/category/brand product hierarchies, or business-line/product/channel organizational structures. If one simply models each series individually, it is difficult to leverage this hierarchical structure; while mixing all series together indiscriminately ignores their individual differences. Hybrid and hierarchical models are designed precisely to address such problems.In real business, time series are rarely "isolated": they often have clear hierarchical structures and shared patterns — for example, store/city/region/national sales hierarchies, SKU/category/brand product hierarchies, or business-line/product/channel organizational structures. If one simply models each series individually, it is difficult to leverage this hierarchical structure; while mixing all series together indiscriminately ignores their individual differences. Hybrid and hierarchical models are designed precisely to address such problems.
One common approach is the global + local model: a shared "global model" learns the common patterns of all series (such as overall trends, holiday effects, seasonality), while introducing local parameters or embedding vectors for each series or each sub-group to capture individual characteristics. This structure avoids the data sparsity problem of training separate models for long-tail series while retaining the ability to finely model popular series.One common approach is the global + local model: a shared "global model" learns the common patterns of all series (such as overall trends, holiday effects, seasonality), while introducing local parameters or embedding vectors for each series or each sub-group to capture individual characteristics. This structure avoids the data sparsity problem of training separate models for long-tail series while retaining the ability to finely model popular series.
Another category is hierarchical time series (hierarchical TS) modeling: explicitly considering hierarchical constraints during the forecasting process (e.g., the sum of sub-level forecasts should be consistent with the upper-level forecast), and through top-down, bottom-up, or mid-level joint optimization, ensuring that forecasts at all levels are consistent in both value and structure. Under deep time series frameworks, this typically manifests as incorporating hierarchical features in input encoding, designing multi-head outputs for different levels, or training with hierarchical loss functions.Another category is hierarchical time series (hierarchical TS) modeling: explicitly considering hierarchical constraints during the forecasting process (e.g., the sum of sub-level forecasts should be consistent with the upper-level forecast), and through top-down, bottom-up, or mid-level joint optimization, ensuring that forecasts at all levels are consistent in both value and structure. Under deep time series frameworks, this typically manifests as incorporating hierarchical features in input encoding, designing multi-head outputs for different levels, or training with hierarchical loss functions.
From a product perspective, this type of hybrid and hierarchical modeling is widely applied in e-commerce sales forecasting platforms, supply-demand/capacity/scheduling forecasting systems, and similar scenarios: the system needs to simultaneously provide forecasts at different granularities such as "single store, single product," "city level," and "national total," and maintain consistency between upper and lower levels during resource planning and KPI decomposition. The flexible structure of deep models allows such constraints to be embedded into the modeling process end-to-end, rather than relying entirely on post-hoc corrections.From a product perspective, this type of hybrid and hierarchical modeling is widely applied in e-commerce sales forecasting platforms, supply-demand/capacity/scheduling forecasting systems, and similar scenarios: the system needs to simultaneously provide forecasts at different granularities such as "single store, single product," "city level," and "national total," and maintain consistency between upper and lower levels during resource planning and KPI decomposition. The flexible structure of deep models allows such constraints to be embedded into the modeling process end-to-end, rather than relying entirely on post-hoc corrections.
In time series scenarios, "forecasting the future" is only part of the problem; another equally critical part is: real-time discovery of anomalies and structural changes. Whether it is equipment operation, business metrics, transaction behavior, or operations monitoring, anomaly detection and change point detection are core capabilities for ensuring system stability and identifying risks and opportunities. Traditionally, statistical threshold methods, EWMA, CUSUM, and others have been widely used; as data dimensionality and complexity increase, various machine learning and deep learning methods (Isolation Forest, One‑Class SVM, AutoEncoder/VAE, time series GAN, GNN + time series models) have also begun to play important roles.In time series scenarios, "forecasting the future" is only part of the problem; another equally critical part is: real-time discovery of anomalies and structural changes. Whether it is equipment operation, business metrics, transaction behavior, or operations monitoring, anomaly detection and change point detection are core capabilities for ensuring system stability and identifying risks and opportunities. Traditionally, statistical threshold methods, EWMA, CUSUM, and others have been widely used; as data dimensionality and complexity increase, various machine learning and deep learning methods (Isolation Forest, One‑Class SVM, AutoEncoder/VAE, time series GAN, GNN + time series models) have also begun to play important roles.
In terms of product form, such capabilities are often embedded in equipment fault early-warning systems, business metric anomaly alerting platforms (e.g., sudden conversion rate drops), security attack and fraud detection systems, and operations AIOps alerting engines, automatically flagging suspicious points and structural changes by monitoring multidimensional time series signals in real time, and integrating with rules, knowledge bases, and manual decision-making processes. Below, we continue to elaborate from three angles: scenarios, principles, and models.In terms of product form, such capabilities are often embedded in equipment fault early-warning systems, business metric anomaly alerting platforms (e.g., sudden conversion rate drops), security attack and fraud detection systems, and operations AIOps alerting engines, automatically flagging suspicious points and structural changes by monitoring multidimensional time series signals in real time, and integrating with rules, knowledge bases, and manual decision-making processes. Below, we continue to elaborate from three angles: scenarios, principles, and models.
Anomaly and change point detection is essentially about finding significant deviations and structural breaks from "normal patterns":Anomaly and change point detection is essentially about finding significant deviations and structural breaks from "normal patterns":
From the perspective of method families, they can be roughly divided into statistical methods, one-class/isolation learning methods, reconstruction-based deep models, and graph + time series combination models:From the perspective of method families, they can be roughly divided into statistical methods, one-class/isolation learning methods, reconstruction-based deep models, and graph + time series combination models:
Below, we elaborate on three directions: point/sequence anomalies, change point detection, and multidimensional and graph-structured approaches.Below, we elaborate on three directions: point/sequence anomalies, change point detection, and multidimensional and graph-structured approaches.
The most intuitive form of anomaly detection is point anomalies: the observed value at a certain time point is far from the historical normal range (e.g., CPU usage suddenly spikes to 100%, transaction amount is abnormally large, sensor reading jumps instantaneously). In traditional methods, the most common practice is to fit a statistical distribution or sliding statistics (mean, variance, quantiles) to historical normal data, set thresholds or control charts (such as EWMA, CUSUM) on this basis, and issue an alert when the current observation exceeds the acceptable interval. The advantages are simple implementation, low computational cost, and easy interpretability, so they remain widely used in a large number of operations monitoring and industrial systems.The most intuitive form of anomaly detection is point anomalies: the observed value at a certain time point is far from the historical normal range (e.g., CPU usage suddenly spikes to 100%, transaction amount is abnormally large, sensor reading jumps instantaneously). In traditional methods, the most common practice is to fit a statistical distribution or sliding statistics (mean, variance, quantiles) to historical normal data, set thresholds or control charts (such as EWMA, CUSUM) on this basis, and issue an alert when the current observation exceeds the acceptable interval. The advantages are simple implementation, low computational cost, and easy interpretability, so they remain widely used in a large number of operations monitoring and industrial systems.
When dimensionality increases or patterns become more complex, one can introduce one-class/isolation learning methods such as Isolation Forest and One‑Class SVM: they learn an aggregated region (or boundary) on "normal samples" and treat points falling outside this region as anomalies. By extracting statistical features (such as window mean, variance, frequency-domain features, etc.) on sliding windows of the series, these methods can also be used to identify local "sequence anomalies" (i.e., behavior deviating from normal patterns over a period of time), suitable for multidimensional metrics and scenarios where the distribution shape is difficult to precisely define.When dimensionality increases or patterns become more complex, one can introduce one-class/isolation learning methods such as Isolation Forest and One‑Class SVM: they learn an aggregated region (or boundary) on "normal samples" and treat points falling outside this region as anomalies. By extracting statistical features (such as window mean, variance, frequency-domain features, etc.) on sliding windows of the series, these methods can also be used to identify local "sequence anomalies" (i.e., behavior deviating from normal patterns over a period of time), suitable for multidimensional metrics and scenarios where the distribution shape is difficult to precisely define.
Under the deep learning framework, methods based on reconstruction error using AutoEncoder / VAE / time series GAN offer more flexible options:Under the deep learning framework, methods based on reconstruction error using AutoEncoder / VAE / time series GAN offer more flexible options:
These methods can adapt to highly nonlinear patterns and complex covariate structures, making them particularly suitable for building unified anomaly detection engines on multidimensional business metrics and complex equipment sensor data.These methods can adapt to highly nonlinear patterns and complex covariate structures, making them particularly suitable for building unified anomaly detection engines on multidimensional business metrics and complex equipment sensor data.
Unlike point anomalies and local anomalies, change point detection focuses on structural breaks in time series: for example, the mean jumps from one level to another, volatility changes, or periodicity and correlation structures adjust. Such changes often correspond to some kind of event or state switch in the real world, such as configuration changes, new policy activation, policy adjustments, production process changes, or market regime switches, and are extremely critical for business diagnosis and causal analysis.Unlike point anomalies and local anomalies, change point detection focuses on structural breaks in time series: for example, the mean jumps from one level to another, volatility changes, or periodicity and correlation structures adjust. Such changes often correspond to some kind of event or state switch in the real world, such as configuration changes, new policy activation, policy adjustments, production process changes, or market regime switches, and are extremely critical for business diagnosis and causal analysis.
Among traditional statistical methods, change point detection often employs techniques such as likelihood ratio tests, CUSUM, and Bayesian Online Change Point Detection (BOCPD):Among traditional statistical methods, change point detection often employs techniques such as likelihood ratio tests, CUSUM, and Bayesian Online Change Point Detection (BOCPD):
In more complex settings, one can combine deep representation learning with segmentation models, treating change point detection as a sequence segmentation problem: use neural networks to extract features, then find segment boundaries in the feature space, or directly train a model to predict the probability that a given time point is a "change point." This is particularly useful for business metrics that exhibit multiple forms of change (not just mean/variance changes) and are difficult to characterize with simple statistical assumptions.In more complex settings, one can combine deep representation learning with segmentation models, treating change point detection as a sequence segmentation problem: use neural networks to extract features, then find segment boundaries in the feature space, or directly train a model to predict the probability that a given time point is a "change point." This is particularly useful for business metrics that exhibit multiple forms of change (not just mean/variance changes) and are difficult to characterize with simple statistical assumptions.
In product systems, change point detection is typically integrated into business metric analysis platforms, A/B experiment analysis systems, and configuration and policy change monitoring tools: when key metrics show structural changes, the system can automatically flag potential change points and associate them with corresponding change events (such as version releases, parameter adjustments, policy implementations), providing clues for subsequent root cause analysis.In product systems, change point detection is typically integrated into business metric analysis platforms, A/B experiment analysis systems, and configuration and policy change monitoring tools: when key metrics show structural changes, the system can automatically flag potential change points and associate them with corresponding change events (such as version releases, parameter adjustments, policy implementations), providing clues for subsequent root cause analysis.
In modern distributed systems and IoT scenarios, we often face multi-point, multidimensional time series with associated topological structures: for example, multiple measurement points in a sensor network, various service metrics in a microservice architecture, or multiple nodes and edges in a power distribution/transportation network. In such cases, performing anomaly detection on each time series individually and in isolation can easily misjudge local fluctuations or ignore overall patterns — true anomalies are often manifestations of "local-global inconsistency" or "topological structure incoordination."In modern distributed systems and IoT scenarios, we often face multi-point, multidimensional time series with associated topological structures: for example, multiple measurement points in a sensor network, various service metrics in a microservice architecture, or multiple nodes and edges in a power distribution/transportation network. In such cases, performing anomaly detection on each time series individually and in isolation can easily misjudge local fluctuations or ignore overall patterns — true anomalies are often manifestations of "local-global inconsistency" or "topological structure incoordination."
To this end, a large number of combination methods of Graph Neural Networks (GNN) + time series models have emerged in recent years:To this end, a large number of combination methods of Graph Neural Networks (GNN) + time series models have emerged in recent years:
This framework is particularly applicable in scenarios such as sensor network monitoring, microservice metric anomaly detection, and spatio-temporal anomaly detection in urban computing: it can distinguish "global changes" (such as an overall system load increase) from "local anomalies" (such as abnormal congestion at a single node), and can also better identify topology-related anomaly patterns (such as link-level problems, regional network failures).This framework is particularly applicable in scenarios such as sensor network monitoring, microservice metric anomaly detection, and spatio-temporal anomaly detection in urban computing: it can distinguish "global changes" (such as an overall system load increase) from "local anomalies" (such as abnormal congestion at a single node), and can also better identify topology-related anomaly patterns (such as link-level problems, regional network failures).
At the engineering level, such methods typically appear as advanced capabilities of operations AIOps alerting systems, security and risk control platforms, and equipment fleet monitoring systems, combining basic statistical monitoring, rule systems, and expert knowledge to provide more intelligent, more context-aware anomaly discovery mechanisms for complex systems.At the engineering level, such methods typically appear as advanced capabilities of operations AIOps alerting systems, security and risk control platforms, and equipment fleet monitoring systems, combining basic statistical monitoring, rule systems, and expert knowledge to provide more intelligent, more context-aware anomaly discovery mechanisms for complex systems.
In many critical business scenarios, modeling only "time" is insufficient: "when" and "where" coexist, and the two are highly coupled. Urban traffic flow is jointly influenced by road network structure and temporal patterns; meteorology and air quality depend on both temporal evolution and geographic proximity and atmospheric flow fields; logistics, bike-sharing, and ride-hailing dispatch need to simultaneously consider the spatio-temporal distribution of demand and road/area structure. Spatio-temporal sequence modeling is precisely the systematic approach for such "time + space" joint modeling problems.In many critical business scenarios, modeling only "time" is insufficient: "when" and "where" coexist, and the two are highly coupled. Urban traffic flow is jointly influenced by road network structure and temporal patterns; meteorology and air quality depend on both temporal evolution and geographic proximity and atmospheric flow fields; logistics, bike-sharing, and ride-hailing dispatch need to simultaneously consider the spatio-temporal distribution of demand and road/area structure. Spatio-temporal sequence modeling is precisely the systematic approach for such "time + space" joint modeling problems.
Compared to pure time series models, spatio-temporal models need to explicitly incorporate spatial dependency structure: traffic flow on adjacent road segments, air quality at nearby monitoring stations, and load and status of connected nodes are usually more correlated than points farther apart. To this end, structures such as Graph Neural Networks (GNN) and Convolutional LSTM (ConvLSTM) are widely used to combine feature learning in both spatial and temporal dimensions. At the product level, such capabilities support a large number of critical applications in urban computing platforms (traffic/pedestrian flow forecasting), meteorological/environmental forecasting systems, logistics route planning, and bike-sharing/ride-hailing dispatch platforms.Compared to pure time series models, spatio-temporal models need to explicitly incorporate spatial dependency structure: traffic flow on adjacent road segments, air quality at nearby monitoring stations, and load and status of connected nodes are usually more correlated than points farther apart. To this end, structures such as Graph Neural Networks (GNN) and Convolutional LSTM (ConvLSTM) are widely used to combine feature learning in both spatial and temporal dimensions. At the product level, such capabilities support a large number of critical applications in urban computing platforms (traffic/pedestrian flow forecasting), meteorological/environmental forecasting systems, logistics route planning, and bike-sharing/ride-hailing dispatch platforms.
The core of spatio-temporal sequence modeling is simultaneously learning spatial correlation and temporal dynamics within a unified framework:The core of spatio-temporal sequence modeling is simultaneously learning spatial correlation and temporal dynamics within a unified framework:
Typical spatio-temporal models mostly adopt the combination form of "GNN + time series model" or "convolution + LSTM":Typical spatio-temporal models mostly adopt the combination form of "GNN + time series model" or "convolution + LSTM":
Below, we elaborate on three directions: spatio-temporal tasks and data representation, GNN + time series models, and convolutional LSTM and spatio-temporal convolution.Below, we elaborate on three directions: spatio-temporal tasks and data representation, GNN + time series models, and convolutional LSTM and spatio-temporal convolution.
Before diving into specific models, the first thing spatio-temporal sequence modeling must address is how to represent spatial structure. Unlike the one-dimensional time axis, spatial structures can be regular grids, irregular graphs, or hybrid forms.Before diving into specific models, the first thing spatio-temporal sequence modeling must address is how to represent spatial structure. Unlike the one-dimensional time axis, spatial structures can be regular grids, irregular graphs, or hybrid forms.
This unified representation of "spatial structure + time series" allows many different scenarios to be modeled as similar problems: given historical spatio-temporal sequences, predict the state of each node or grid cell at several future time steps. Subsequent model designs (whether GNN + time series models, or ConvLSTM) are all developed from this unified perspective.This unified representation of "spatial structure + time series" allows many different scenarios to be modeled as similar problems: given historical spatio-temporal sequences, predict the state of each node or grid cell at several future time steps. Subsequent model designs (whether GNN + time series models, or ConvLSTM) are all developed from this unified perspective.
At the product level, this layer of abstraction is often encapsulated in the data layer and modeling layer of urban computing platforms, meteorological/environmental forecasting systems, and route planning and dispatch platforms: business stakeholders only need to know "we are predicting future traffic/demand on the road network/grid," while the underlying data representation and spatio-temporal fusion are handled uniformly by the modeling framework.At the product level, this layer of abstraction is often encapsulated in the data layer and modeling layer of urban computing platforms, meteorological/environmental forecasting systems, and route planning and dispatch platforms: business stakeholders only need to know "we are predicting future traffic/demand on the road network/grid," while the underlying data representation and spatio-temporal fusion are handled uniformly by the modeling framework.
For modeling spatio-temporal sequences on graph structures, the current mainstream approach is the combination of Graph Neural Networks (GNN) + time series models. Representative models include ST‑GCN, DCRNN, Graph WaveNet, ST‑Transformer, and their common characteristics are:For modeling spatio-temporal sequences on graph structures, the current mainstream approach is the combination of Graph Neural Networks (GNN) + time series models. Representative models include ST‑GCN, DCRNN, Graph WaveNet, ST‑Transformer, and their common characteristics are:
For example, DCRNN (Diffusion Convolutional RNN) combines graph convolution with gated recurrent units, using diffusion convolution to simulate the propagation of information on road networks, and then capturing temporal dynamics through RNN, making it very suitable for tasks such as traffic flow forecasting. Graph WaveNet introduces adaptive graph structure learning and multi-scale modeling on top of graph convolution and temporal convolution, improving adaptability to complex road networks and irregular topologies. Models such as ST‑Transformer introduce self-attention mechanisms into spatio-temporal modeling, simultaneously considering correlations between different temporal and spatial positions through spatio-temporal attention modules.For example, DCRNN (Diffusion Convolutional RNN) combines graph convolution with gated recurrent units, using diffusion convolution to simulate the propagation of information on road networks, and then capturing temporal dynamics through RNN, making it very suitable for tasks such as traffic flow forecasting. Graph WaveNet introduces adaptive graph structure learning and multi-scale modeling on top of graph convolution and temporal convolution, improving adaptability to complex road networks and irregular topologies. Models such as ST‑Transformer introduce self-attention mechanisms into spatio-temporal modeling, simultaneously considering correlations between different temporal and spatial positions through spatio-temporal attention modules.
In real systems, this category of GNN + time series models is widely deployed in products such as urban traffic and pedestrian flow forecasting platforms, shared mobility dispatch systems, and complex IoT network monitoring. They typically serve as one of the core forecasting engines, forming a closed loop together with rule systems, simulation models, and business strategies, enabling dispatch and planning to consider both global structure and respond to local changes.In real systems, this category of GNN + time series models is widely deployed in products such as urban traffic and pedestrian flow forecasting platforms, shared mobility dispatch systems, and complex IoT network monitoring. They typically serve as one of the core forecasting engines, forming a closed loop together with rule systems, simulation models, and business strategies, enabling dispatch and planning to consider both global structure and respond to local changes.
Another important route is spatio-temporal modeling based on Convolutional LSTM (ConvLSTM) and its variants. Unlike standard LSTM, which passes one-dimensional vectors between time steps, ConvLSTM uses convolution operators in the gating structure, so that both the hidden state and input remain as multi-dimensional tensors (such as feature maps on a spatial grid). In this way, each time step's state update includes both temporal recurrence and local convolution aggregation in the spatial dimension, achieving natural modeling of spatio-temporal local patterns.Another important route is spatio-temporal modeling based on Convolutional LSTM (ConvLSTM) and its variants. Unlike standard LSTM, which passes one-dimensional vectors between time steps, ConvLSTM uses convolution operators in the gating structure, so that both the hidden state and input remain as multi-dimensional tensors (such as feature maps on a spatial grid). In this way, each time step's state update includes both temporal recurrence and local convolution aggregation in the spatial dimension, achieving natural modeling of spatio-temporal local patterns.
Building on this, improved models such as Conv‑TT‑LSTM attempt to enhance model expressiveness and efficiency through mechanisms such as tensor decomposition, parameter sharing, and multi-scale convolution, adapting to larger-scale, more complex spatio-temporal data. For example, in meteorological forecasting, one can stack multiple layers of ConvLSTM to perform spatio-temporal recurrence on multi-channel meteorological element maps (temperature, humidity, wind direction, etc.), predicting the spatial distribution for the next few hours or days from several historical frames; in traffic and environmental monitoring, one can also map road networks or monitoring points onto regular grids and use models such as ConvLSTM for forecasting.Building on this, improved models such as Conv‑TT‑LSTM attempt to enhance model expressiveness and efficiency through mechanisms such as tensor decomposition, parameter sharing, and multi-scale convolution, adapting to larger-scale, more complex spatio-temporal data. For example, in meteorological forecasting, one can stack multiple layers of ConvLSTM to perform spatio-temporal recurrence on multi-channel meteorological element maps (temperature, humidity, wind direction, etc.), predicting the spatial distribution for the next few hours or days from several historical frames; in traffic and environmental monitoring, one can also map road networks or monitoring points onto regular grids and use models such as ConvLSTM for forecasting.
Compared to GNN + time series models, the ConvLSTM family is more commonly used in scenarios with regular grid structures and pronounced local spatial smoothness, such as meteorological radar echo forecasting, air quality grid forecasting, and video frame-level prediction. Its advantages lie in relatively direct implementation, ease of leveraging existing convolutional network infrastructure for acceleration and deployment, and ease of collaboration with visual models such as CNN/ViT, for example, combining convolutional features and temporal recurrence in remote sensing image spatio-temporal modeling.Compared to GNN + time series models, the ConvLSTM family is more commonly used in scenarios with regular grid structures and pronounced local spatial smoothness, such as meteorological radar echo forecasting, air quality grid forecasting, and video frame-level prediction. Its advantages lie in relatively direct implementation, ease of leveraging existing convolutional network infrastructure for acceleration and deployment, and ease of collaboration with visual models such as CNN/ViT, for example, combining convolutional features and temporal recurrence in remote sensing image spatio-temporal modeling.
In terms of product form, models in this direction are mostly used in meteorological/environmental forecasting systems, remote sensing spatio-temporal analysis platforms, and video and image spatio-temporal forecasting, often exposing their capabilities upstream in the form of "future spatio-temporal scenario forecast maps," serving as important input for business decision-making and visual analysis.In terms of product form, models in this direction are mostly used in meteorological/environmental forecasting systems, remote sensing spatio-temporal analysis platforms, and video and image spatio-temporal forecasting, often exposing their capabilities upstream in the form of "future spatio-temporal scenario forecast maps," serving as important input for business decision-making and visual analysis.
In the previous capability layers such as vision and language, models mostly operated in a "passive answering" mode — receiving input and producing output. In many real-world business scenarios, however, what we need is an agent that can proactively plan, invoke external tools, and orchestrate workflows: it must not only see/read/hear, but also "decide what to do next" on its own — such as looking up information, running code, reading/writing files, calling internal systems, then integrating, interpreting, and feeding the results back to the user.In the previous capability layers such as vision and language, models mostly operated in a "passive answering" mode — receiving input and producing output. In many real-world business scenarios, however, what we need is an agent that can proactively plan, invoke external tools, and orchestrate workflows: it must not only see/read/hear, but also "decide what to do next" on its own — such as looking up information, running code, reading/writing files, calling internal systems, then integrating, interpreting, and feeding the results back to the user.
This layer can be understood as the critical glue that "turns a foundation model into an actionable system": through structured tool-calling interfaces, workflow orchestration, multi-agent collaboration, and human-in-the-loop mechanisms, it extends an LLM from a powerful "cognitive kernel" into a "digital employee" capable of completing end-to-end tasks.This layer can be understood as the critical glue that "turns a foundation model into an actionable system": through structured tool-calling interfaces, workflow orchestration, multi-agent collaboration, and human-in-the-loop mechanisms, it extends an LLM from a powerful "cognitive kernel" into a "digital employee" capable of completing end-to-end tasks.
In the plain-text era of read-only, say-only interactions, an LLM was more like a "super conversationalist": it could understand questions, give advice, write code, and outline plans, but all the "truly executable" work — querying databases, running scripts, generating files, invoking cloud services — still had to be done manually by humans. The emergence of Tool Calling / Function Calling allows models, for the first time, to "take action" within a safety boundary: they automatically generate structured parameters from natural language to invoke external capabilities such as search engines, databases, computation engines, and image/audio/video generation services, and then organize and return the execution results, thereby forming a closed loop of "understanding → decision → execution."In the plain-text era of read-only, say-only interactions, an LLM was more like a "super conversationalist": it could understand questions, give advice, write code, and outline plans, but all the "truly executable" work — querying databases, running scripts, generating files, invoking cloud services — still had to be done manually by humans. The emergence of Tool Calling / Function Calling allows models, for the first time, to "take action" within a safety boundary: they automatically generate structured parameters from natural language to invoke external capabilities such as search engines, databases, computation engines, and image/audio/video generation services, and then organize and return the execution results, thereby forming a closed loop of "understanding → decision → execution."
From a product perspective, tool calling is the "chassis capability" of virtually all agent systems: the OpenAI Assistants API, LangChain, LlamaIndex, AutoGen, and agent platforms from various cloud providers are essentially building a runtime layer on top of LLMs centered around how to define tools, how to let the model correctly select tools, and how to handle errors and retries. Below, we examine this layer from three angles — scenarios, principles, and models — and then expand on "tool-calling interface design," "tool selection and strategy," and "typical tool types" in the subsequent subsections.From a product perspective, tool calling is the "chassis capability" of virtually all agent systems: the OpenAI Assistants API, LangChain, LlamaIndex, AutoGen, and agent platforms from various cloud providers are essentially building a runtime layer on top of LLMs centered around how to define tools, how to let the model correctly select tools, and how to handle errors and retries. Below, we examine this layer from three angles — scenarios, principles, and models — and then expand on "tool-calling interface design," "tool selection and strategy," and "typical tool types" in the subsequent subsections.
The core of tool calling is: driving structured function calls with natural language.The core of tool calling is: driving structured function calls with natural language.
The models and frameworks supporting this capability fall into three main categories:The models and frameworks supporting this capability fall into three main categories:
A usable tool-calling system first requires a clear, standardized, LLM-friendly "tool interface layer." This layer is responsible for wrapping external-world APIs, scripts, and services into "functions" that the model can understand and safely invoke, enabling the model to "verbalize" the tool and its parameters it wishes to call as if writing pseudocode.A usable tool-calling system first requires a clear, standardized, LLM-friendly "tool interface layer." This layer is responsible for wrapping external-world APIs, scripts, and services into "functions" that the model can understand and safely invoke, enabling the model to "verbalize" the tool and its parameters it wishes to call as if writing pseudocode.
At the interface layer, each tool is typically defined using a structure similar to JSON Schema or a function signature: including the name, description, parameter fields (properties), types (string / number / boolean / array / object), whether they are required, and allowed value ranges or enumerations.At the interface layer, each tool is typically defined using a structure similar to JSON Schema or a function signature: including the name, description, parameter fields (properties), types (string / number / boolean / array / object), whether they are required, and allowed value ranges or enumerations.
This information is used on one hand to drive type checking in the frontend/SDK, and on the other hand is fed directly to the LLM to help the model "learn" how to correctly fill in parameters. The clearer the descriptions and the more reasonable the constraints, the more standardized the model's generated calls and the lower the error rate.This information is used on one hand to drive type checking in the frontend/SDK, and on the other hand is fed directly to the LLM to help the model "learn" how to correctly fill in parameters. The clearer the descriptions and the more reasonable the constraints, the more standardized the model's generated calls and the lower the error rate.
When a user makes a request like "look up Q3 2024 revenue and draw a bar chart broken down by region," the model must first reason that this requires at least a "report query tool" (to access data) and possibly a "chart generation tool" (to draw the chart). For each tool, it must extract and map structured parameters from the raw language — such as time range (start_date/end_date), dimension (region), metric (revenue), chart type (bar), output format, etc. — and then output them as JSON to be handed off to the runtime.When a user makes a request like "look up Q3 2024 revenue and draw a bar chart broken down by region," the model must first reason that this requires at least a "report query tool" (to access data) and possibly a "chart generation tool" (to draw the chart). For each tool, it must extract and map structured parameters from the raw language — such as time range (start_date/end_date), dimension (region), metric (revenue), chart type (bar), output format, etc. — and then output them as JSON to be handed off to the runtime.
In this process, the model is essentially performing integrated reasoning of "natural language → task planning → parameter extraction/filling," making the natural-language prompts in tool descriptions, parameter examples, and few-shot samples critically important.In this process, the model is essentially performing integrated reasoning of "natural language → task planning → parameter extraction/filling," making the natural-language prompts in tool descriptions, parameter examples, and few-shot samples critically important.
After receiving the JSON call produced by the model, the runtime first performs parameter validation and security checks before actually invoking the backend API or program. Once execution is complete, the result is wrapped as a structured object (such as a query result table, file URL, media resource ID, etc.) and returned to the model.After receiving the JSON call produced by the model, the runtime first performs parameter validation and security checks before actually invoking the backend API or program. Once execution is complete, the result is wrapped as a structured object (such as a query result table, file URL, media resource ID, etc.) and returned to the model.
The model then transforms these raw results into a user-readable explanation or performs further processing, such as summarizing reports, generating natural-language analysis, or embedding chart annotations. For the model, tool results are merely intermediate information — it is still responsible for "understanding the results + explaining the results."The model then transforms these raw results into a user-readable explanation or performs further processing, such as summarizing reports, generating natural-language analysis, or embedding chart annotations. For the model, tool results are merely intermediate information — it is still responsible for "understanding the results + explaining the results."
When there is only one tool in the system, "whether to use a tool" is the only question. But in real-world agent applications, there are often dozens or even hundreds of tools: retrieval from different data sources, business APIs from different departments, generation/analysis capabilities from different technical domains. This introduces a new challenge: how the model makes reasonable selection and orchestration in a multi-tool environment.When there is only one tool in the system, "whether to use a tool" is the only question. But in real-world agent applications, there are often dozens or even hundreds of tools: retrieval from different data sources, business APIs from different departments, generation/analysis capabilities from different technical domains. This introduces a new challenge: how the model makes reasonable selection and orchestration in a multi-tool environment.
First, the model must determine "whether the current request requires a tool call" and "which tool (or tools) to call." This is typically achieved by listing available tool descriptions in the system prompt and providing typical examples so the model learns to select the appropriate tool based on user intent.First, the model must determine "whether the current request requires a tool call" and "which tool (or tools) to call." This is typically achieved by listing available tool descriptions in the system prompt and providing typical examples so the model learns to select the appropriate tool based on user intent.
For scenarios with a large number of tools and high description similarity, many frameworks introduce a "tool router" (such as vector-retrieval-based or rule-based pre-filtering) that first narrows down a large list to several candidate tools, which are then exposed to the LLM for selection, thereby reducing the model's burden and the probability of misselection.For scenarios with a large number of tools and high description similarity, many frameworks introduce a "tool router" (such as vector-retrieval-based or rule-based pre-filtering) that first narrows down a large list to several candidate tools, which are then exposed to the LLM for selection, thereby reducing the model's burden and the probability of misselection.
Complex tasks often require multiple tools working together. For example, "research the major listed companies in a certain industry and generate a report with financial comparison charts" may involve a search engine, a financial report database, a computation engine, a chart generation tool, and a document export tool.Complex tasks often require multiple tools working together. For example, "research the major listed companies in a certain industry and generate a report with financial comparison charts" may involve a search engine, a financial report database, a computation engine, a chart generation tool, and a document export tool.
In such cases, the model needs to perform lightweight task planning: which tool to use first to obtain a list, then query detailed information for each item on the list one by one, then merge the data, perform calculations and visualization, and finally invoke the export tool to generate the report. Typical practices include ReAct/Planner-Executor approaches, where the model progressively completes combined tool calls in a cycle of "thinking (Plan) — calling (Act) — reflecting (Reflect)."In such cases, the model needs to perform lightweight task planning: which tool to use first to obtain a list, then query detailed information for each item on the list one by one, then merge the data, perform calculations and visualization, and finally invoke the export tool to generate the report. Typical practices include ReAct/Planner-Executor approaches, where the model progressively completes combined tool calls in a cycle of "thinking (Plan) — calling (Act) — reflecting (Reflect)."
Different types of tools provide agent systems with different dimensions of "external brains." From an engineering practice perspective, the following categories of tools are almost "standard equipment" for all complex applications.Different types of tools provide agent systems with different dimensions of "external brains." From an engineering practice perspective, the following categories of tools are almost "standard equipment" for all complex applications.
Retrieval tools are responsible for extending "memory" to the external world:Retrieval tools are responsible for extending "memory" to the external world:
In RAG scenarios, the LLM pulls context relevant to the user's question via retrieval tools and then performs reasoning and generation on top of that context, significantly improving the timeliness and accuracy of answers.In RAG scenarios, the LLM pulls context relevant to the user's question via retrieval tools and then performs reasoning and generation on top of that context, significantly improving the timeliness and accuracy of answers.
Code execution tools (such as Python/JS sandboxes, notebook executors) allow LLMs to "write a piece of code and run it immediately" to solve complex computation, data processing, numerical simulation, visualization, and other problems.Code execution tools (such as Python/JS sandboxes, notebook executors) allow LLMs to "write a piece of code and run it immediately" to solve complex computation, data processing, numerical simulation, visualization, and other problems.
The model is responsible for producing code and input parameters, while the execution environment is responsible for security isolation, resource limiting, and result collection. These tools are critical in scenarios such as data analysis, quantitative research, automated reporting, scientific computing, and agent self-verification (where the model generates an answer and then verifies it with code).The model is responsible for producing code and input parameters, while the execution environment is responsible for security isolation, resource limiting, and result collection. These tools are critical in scenarios such as data analysis, quantitative research, automated reporting, scientific computing, and agent self-verification (where the model generates an answer and then verifies it with code).
File read/write tools are responsible for bringing external file systems and data sources into the agent's field of view: reading PDF/Word/Excel, accessing database tables, calling internal business APIs, etc. The model obtains real business data through these tools and then performs summarization, comparison, and report generation.File read/write tools are responsible for bringing external file systems and data sources into the agent's field of view: reading PDF/Word/Excel, accessing database tables, calling internal business APIs, etc. The model obtains real business data through these tools and then performs summarization, comparison, and report generation.
Accompanying these are file writing and management tools: persisting generated reports, charts, PPTs, code, etc., and returning links or IDs for convenient subsequent access and integration by users.Accompanying these are file writing and management tools: persisting generated reports, charts, PPTs, code, etc., and returning links or IDs for convenient subsequent access and integration by users.
Media generation tools add "creative" and "design" arms to agents:Media generation tools add "creative" and "design" arms to agents:
In content production, marketing design, education and training, gaming, and multimedia applications, these tools bring "from idea to finished product" closer to an automated pipeline.In content production, marketing design, education and training, gaming, and multimedia applications, these tools bring "from idea to finished product" closer to an automated pipeline.
Taken together, tool calling and execution extend LLMs from "language models" to "general-purpose controllers with action interfaces": the model understands requirements and the environment through language, performs real operations through tools, and continuously adjusts its strategy through feedback. Combined with appropriate workflow orchestration and multi-agent collaboration (see 7.2), this forms the foundational architecture for the next generation of intelligent applications.Taken together, tool calling and execution extend LLMs from "language models" to "general-purpose controllers with action interfaces": the model understands requirements and the environment through language, performs real operations through tools, and continuously adjusts its strategy through feedback. Combined with appropriate workflow orchestration and multi-agent collaboration (see 7.2), this forms the foundational architecture for the next generation of intelligent applications.
With tool-calling capabilities, an LLM is no longer just a "question answerer" but can become an "execution unit" oriented toward specific tasks. However, real-world business is often far more complex than a single conversation: a complete litigation analysis, a market research project, an A/B experiment configuration, or an end-to-end ops handling process typically requires multi-step operations, multiple tools, and even long-term participation from multiple roles. At this point, the single LLM + tools pattern begins to strain, necessitating further workflow orchestration and multi-agent collaboration.With tool-calling capabilities, an LLM is no longer just a "question answerer" but can become an "execution unit" oriented toward specific tasks. However, real-world business is often far more complex than a single conversation: a complete litigation analysis, a market research project, an A/B experiment configuration, or an end-to-end ops handling process typically requires multi-step operations, multiple tools, and even long-term participation from multiple roles. At this point, the single LLM + tools pattern begins to strain, necessitating further workflow orchestration and multi-agent collaboration.
From a systems perspective, this layer's responsibility is: abstracting a complex, multi-step, multi-stakeholder business process into a workflow graph that LLMs can understand and manipulate, then scheduling one or more agents on this graph, coordinating with human intervention, to jointly complete the task. Typical implementations include Planner-Executor agent architectures, agents with reflection/self-correction capabilities, and graph-based Workflow Orchestrators; corresponding product forms include various automated report generation and operations automation platforms, low-code workflow + LLM integration, complex business process robots, automated ops systems, and so on.From a systems perspective, this layer's responsibility is: abstracting a complex, multi-step, multi-stakeholder business process into a workflow graph that LLMs can understand and manipulate, then scheduling one or more agents on this graph, coordinating with human intervention, to jointly complete the task. Typical implementations include Planner-Executor agent architectures, agents with reflection/self-correction capabilities, and graph-based Workflow Orchestrators; corresponding product forms include various automated report generation and operations automation platforms, low-code workflow + LLM integration, complex business process robots, automated ops systems, and so on.
The core of workflow and multi-agent collaboration is adding a layer of structured control and state management on top of LLMs:The core of workflow and multi-agent collaboration is adding a layer of structured control and state management on top of LLMs:
The main technical directions supporting this layer include:The main technical directions supporting this layer include:
What users give to agents is typically a highly compressed natural-language requirement, such as "do a market research report on the new energy vehicle industry and output a PPT," which actually encompasses numerous steps including retrieval, filtering, analysis, visualization, layout, and multiple rounds of revision. How to automatically build a clear, executable workflow from this single sentence is the first step in workflow orchestration.What users give to agents is typically a highly compressed natural-language requirement, such as "do a market research report on the new energy vehicle industry and output a PPT," which actually encompasses numerous steps including retrieval, filtering, analysis, visualization, layout, and multiple rounds of revision. How to automatically build a clear, executable workflow from this single sentence is the first step in workflow orchestration.
A Planner-type agent must first "unfold" the requirement: combine built-in templates, historical cases, and the tool inventory to identify key phases (such as information collection, data analysis, structural design, content writing, review and export) and further refine them into executable subtasks (such as "retrieve 5 authoritative industry reports from the past year," "pull the last 3 years of sales data broken down by vehicle model," "generate 3 comparison charts," etc.).A Planner-type agent must first "unfold" the requirement: combine built-in templates, historical cases, and the tool inventory to identify key phases (such as information collection, data analysis, structural design, content writing, review and export) and further refine them into executable subtasks (such as "retrieve 5 authoritative industry reports from the past year," "pull the last 3 years of sales data broken down by vehicle model," "generate 3 comparison charts," etc.).
The dependencies and scheduling logic among these subtasks are explicitly represented as a graph or state machine: which ones can run in parallel, which must execute sequentially, at which nodes human confirmation is needed, and under what conditions rollback or retry is required.The dependencies and scheduling logic among these subtasks are explicitly represented as a graph or state machine: which ones can run in parallel, which must execute sequentially, at which nodes human confirmation is needed, and under what conditions rollback or retry is required.
Real-world processes are often not linear pipelines but contain conditional branches (e.g., "if not enough high-quality reports can be retrieved, switch keywords or data sources"), loops (e.g., "continuously attempt rewriting and compression until the report length meets the limit"), and exception paths (e.g., "if a data source is unreachable, switch to an alternative source or use an estimation method").Real-world processes are often not linear pipelines but contain conditional branches (e.g., "if not enough high-quality reports can be retrieved, switch keywords or data sources"), loops (e.g., "continuously attempt rewriting and compression until the report length meets the limit"), and exception paths (e.g., "if a data source is unreachable, switch to an alternative source or use an estimation method").
This requires the workflow orchestration layer to be able to express control-flow semantics such as if/else, while/for, and try/catch on the graph structure, and to allow the Planner Agent or upper-level orchestrator to make decisions based on real-time results during execution, rather than merely planning all steps upfront in a one-shot manner.This requires the workflow orchestration layer to be able to express control-flow semantics such as if/else, while/for, and try/catch on the graph structure, and to allow the Planner Agent or upper-level orchestrator to make decisions based on real-time results during execution, rather than merely planning all steps upfront in a one-shot manner.
Task decomposition and planning are tightly connected with the tool calling discussed in 7.1: when generating subtasks, the Planner often simultaneously specifies "which tools/agents this task requires" and "the input/output format of this node," laying the groundwork for subsequent automatic parameter filling and tool execution.Task decomposition and planning are tightly connected with the tool calling discussed in 7.1: when generating subtasks, the Planner often simultaneously specifies "which tools/agents this task requires" and "the input/output format of this node," laying the groundwork for subsequent automatic parameter filling and tool execution.
Some systems adopt an explicit two-phase "Plan + Execute" approach: the Planner first outputs a machine-readable plan (such as a JSON workflow description), and then the Executor strictly invokes tools and agents according to the plan. Other systems adopt a ReAct style, weaving "thinking–tool calling–observation–rethinking" into the same conversation to achieve more flexible adaptive execution.Some systems adopt an explicit two-phase "Plan + Execute" approach: the Planner first outputs a machine-readable plan (such as a JSON workflow description), and then the Executor strictly invokes tools and agents according to the plan. Other systems adopt a ReAct style, weaving "thinking–tool calling–observation–rethinking" into the same conversation to achieve more flexible adaptive execution.
A single large model is certainly powerful, but in complex business scenarios, different domains often require different knowledge structures, style preferences, and safety strategies. The idea behind multi-agent collaboration is to decompose a single "big and comprehensive" intelligence into multiple "specialized and refined" roles: someone is responsible for planning, someone for execution, someone for review, and someone for domain-specific professional judgment, forming a virtual team composed of agents + tools + humans.A single large model is certainly powerful, but in complex business scenarios, different domains often require different knowledge structures, style preferences, and safety strategies. The idea behind multi-agent collaboration is to decompose a single "big and comprehensive" intelligence into multiple "specialized and refined" roles: someone is responsible for planning, someone for execution, someone for review, and someone for domain-specific professional judgment, forming a virtual team composed of agents + tools + humans.
In a typical multi-agent workflow, common roles include:In a typical multi-agent workflow, common roles include:
For highly specialized domains such as law, finance, engineering, and operations, domain-expert agents can be further subdivided: such as "legal advisor agent," "investment research analyst agent," "cloud-native ops agent," "advertising optimization agent," etc.For highly specialized domains such as law, finance, engineering, and operations, domain-expert agents can be further subdivided: such as "legal advisor agent," "investment research analyst agent," "cloud-native ops agent," "advertising optimization agent," etc.
They can participate in project-based collaboration based on domain-specific knowledge bases, tools, and even specially fine-tuned models: for example, in investment and financing materials, the technical agent handles the technical feasibility section, the financial agent handles the financial model and valuation, the legal agent handles compliance and risk disclosure, the operations agent handles market and growth strategy, and a master control agent consolidates and unifies the style.They can participate in project-based collaboration based on domain-specific knowledge bases, tools, and even specially fine-tuned models: for example, in investment and financing materials, the technical agent handles the technical feasibility section, the financial agent handles the financial model and valuation, the legal agent handles compliance and risk disclosure, the operations agent handles market and growth strategy, and a master control agent consolidates and unifies the style.
The key to multi-agent collaboration also lies in "who talks to whom and when." The system needs a message routing and coordination mechanism to:The key to multi-agent collaboration also lies in "who talks to whom and when." The system needs a message routing and coordination mechanism to:
These capabilities are typically provided by an upper-level orchestrator or "management agent," while frameworks such as LangChain and AutoGen provide infrastructure at the engineering level for conversation routing, multi-agent sessions, role configuration, and more.These capabilities are typically provided by an upper-level orchestrator or "management agent," while frameworks such as LangChain and AutoGen provide infrastructure at the engineering level for conversation routing, multi-agent sessions, role configuration, and more.
No matter how intelligent workflows and multi-agent collaboration become, real-world business still cannot completely dispense with human judgment, especially in high-risk, high-cost, high-sensitivity scenarios such as legal compliance, financial decision-making, medical advice, large-scale production changes, and public opinion response. The design of Human-in-the-loop is precisely about finding a balance between automation and controllability: automate what should be automated, and definitely pause for a human to review what requires manual confirmation.No matter how intelligent workflows and multi-agent collaboration become, real-world business still cannot completely dispense with human judgment, especially in high-risk, high-cost, high-sensitivity scenarios such as legal compliance, financial decision-making, medical advice, large-scale production changes, and public opinion response. The design of Human-in-the-loop is precisely about finding a balance between automation and controllability: automate what should be automated, and definitely pause for a human to review what requires manual confirmation.
In the workflow graph, several "manual approval/confirmation nodes" are typically explicitly marked:In the workflow graph, several "manual approval/confirmation nodes" are typically explicitly marked:
The Orchestrator pauses automatic execution at these nodes, sends intermediate results to the corresponding human roles, and continues the subsequent flow only after receiving feedback.The Orchestrator pauses automatic execution at these nodes, sends intermediate results to the corresponding human roles, and continues the subsequent flow only after receiving feedback.
Humans don't just "press approve or reject" at a certain moment; more importantly, the content of their feedback can be absorbed by the system:Humans don't just "press approve or reject" at a certain moment; more importantly, the content of their feedback can be absorbed by the system:
Finally, Human-in-the-loop also requires a clear set of risk-grading and observability mechanisms:Finally, Human-in-the-loop also requires a clear set of risk-grading and observability mechanisms:
These capabilities not only improve the system's acceptability within the enterprise but also provide a foundation for subsequent compliance audits and responsibility attribution.These capabilities not only improve the system's acceptability within the enterprise but also provide a foundation for subsequent compliance audits and responsibility attribution.
Taken together, tool calling and execution (7.1) addresses the problem of "single-step action," while workflow orchestration and multi-agent collaboration (7.2) attempts to answer "how to string many steps together so that different roles can collaborate long-term and operate in a controllable manner." The combination of both, along with Human-in-the-loop and sound engineering practices, forms the next-generation intelligent application foundation for real-world business scenarios.Taken together, tool calling and execution (7.1) addresses the problem of "single-step action," while workflow orchestration and multi-agent collaboration (7.2) attempts to answer "how to string many steps together so that different roles can collaborate long-term and operate in a controllable manner." The combination of both, along with Human-in-the-loop and sound engineering practices, forms the next-generation intelligent application foundation for real-world business scenarios.
In the preceding vision and understanding layers, models primarily rely on "knowledge learned within their own parameters" to understand and generate content. In real business settings, however, many problems cannot be solved by "memory" alone: internal company policies change daily, regulations and industry standards are continuously updated, and a particular customer's history exists only in internal databases. In these cases, relying solely on what the model has "memorized" is far from sufficient — what matters more is whether the model can efficiently retrieve and reason over external knowledge bases, structured data, and knowledge graphs.In the preceding vision and understanding layers, models primarily rely on "knowledge learned within their own parameters" to understand and generate content. In real business settings, however, many problems cannot be solved by "memory" alone: internal company policies change daily, regulations and industry standards are continuously updated, and a particular customer's history exists only in internal databases. In these cases, relying solely on what the model has "memorized" is far from sufficient — what matters more is whether the model can efficiently retrieve and reason over external knowledge bases, structured data, and knowledge graphs.
Think of this layer as adding an "external brain that can look up references and query databases" on top of the model's capabilities. When a user poses a question, the system no longer generates an answer directly; instead, it first goes to the appropriate data sources to "look things up": document repositories, databases, search engines, knowledge graphs, logs, and business systems... and then has the model produce an answer or decision based on the genuinely retrieved content. This not only significantly improves accuracy and timeliness but also substantially enhances explainability and compliance (for example, citing sources and preserving executed SQL records).Think of this layer as adding an "external brain that can look up references and query databases" on top of the model's capabilities. When a user poses a question, the system no longer generates an answer directly; instead, it first goes to the appropriate data sources to "look things up": document repositories, databases, search engines, knowledge graphs, logs, and business systems... and then has the model produce an answer or decision based on the genuinely retrieved content. This not only significantly improves accuracy and timeliness but also substantially enhances explainability and compliance (for example, citing sources and preserving executed SQL records).
Within this layer, common capabilities broadly fall into two directions: one is Retrieval-Augmented Generation (RAG), primarily oriented toward "natural language Q&A + document/knowledge base retrieval"; the other is Structured Data & Knowledge Graphs (Structured Data & KG), responsible for more precise, controllable access and reasoning over databases, graph databases, and domain knowledge platforms. These are detailed below.Within this layer, common capabilities broadly fall into two directions: one is Retrieval-Augmented Generation (RAG), primarily oriented toward "natural language Q&A + document/knowledge base retrieval"; the other is Structured Data & Knowledge Graphs (Structured Data & KG), responsible for more precise, controllable access and reasoning over databases, graph databases, and domain knowledge platforms. These are detailed below.
RAG (Retrieval‑Augmented Generation) can be thought of as an "LLM that knows how to look things up." Unlike relying purely on a model's internal parameters, RAG, before answering each question, first retrieves from an external knowledge base, finds the most relevant document fragments (chunks), and then feeds those retrieved pieces as "context" to the LLM so that it generates an answer based on "having seen the material." For enterprise knowledge base Q&A, industry report search, legal/medical/financial professional Q&A, internal document search bots, and similar scenarios, RAG has become the default paradigm.RAG (Retrieval‑Augmented Generation) can be thought of as an "LLM that knows how to look things up." Unlike relying purely on a model's internal parameters, RAG, before answering each question, first retrieves from an external knowledge base, finds the most relevant document fragments (chunks), and then feeds those retrieved pieces as "context" to the LLM so that it generates an answer based on "having seen the material." For enterprise knowledge base Q&A, industry report search, legal/medical/financial professional Q&A, internal document search bots, and similar scenarios, RAG has become the default paradigm.
Architecturally, a typical RAG system can be broken down into three layers: the indexing layer, the retrieval layer, and the generation layer. The first two are primarily about "retrieving accurately," while the last is responsible for "articulating clearly." The following sections expand on these three layers, with further refinement of core design and practices in the subsections.Architecturally, a typical RAG system can be broken down into three layers: the indexing layer, the retrieval layer, and the generation layer. The first two are primarily about "retrieving accurately," while the last is responsible for "articulating clearly." The following sections expand on these three layers, with further refinement of core design and practices in the subsections.
The core idea of RAG is "store knowledge externally, delegate reasoning to the model":The core idea of RAG is "store knowledge externally, delegate reasoning to the model":
A typical RAG system is often a model composition architecture:A typical RAG system is often a model composition architecture:
In any RAG system, index construction is foundational. Without high‑quality indexing, even the most powerful LLM downstream is like "a skilled cook without ingredients." The goal of index construction is to transform messy document resources into "retrievable, maintainable, and scalable knowledge assets."In any RAG system, index construction is foundational. Without high‑quality indexing, even the most powerful LLM downstream is like "a skilled cook without ingredients." The goal of index construction is to transform messy document resources into "retrievable, maintainable, and scalable knowledge assets."
From a process perspective, typical index construction involves the following key steps:From a process perspective, typical index construction involves the following key steps:
Documents are often long PDFs, PPTs, Word files, or web pages. Vectorizing an entire document at once both tends to cause "dilution" (one document spans multiple topics) and hinders efficient retrieval. Therefore, you need to:Documents are often long PDFs, PPTs, Word files, or web pages. Vectorizing an entire document at once both tends to cause "dilution" (one document spans multiple topics) and hinders efficient retrieval. Therefore, you need to:
Based on the chunks, generate semantic vectors for each document chunk:Based on the chunks, generate semantic vectors for each document chunk:
Semantic vectors alone cannot meet complex filtering needs; a meta‑information index is usually also required:Semantic vectors alone cannot meet complex filtering needs; a meta‑information index is usually also required:
Once index construction is complete, when a user submits a query, the system enters the retrieval and re‑ranking phase. The key here is not just "finding some relevant documents," but finding an evidence set that is both relevant, sufficiently comprehensive, and supportive of reasoning.Once index construction is complete, when a user submits a query, the system enters the retrieval and re‑ranking phase. The key here is not just "finding some relevant documents," but finding an evidence set that is both relevant, sufficiently comprehensive, and supportive of reasoning.
Pure vector retrieval excels at capturing semantic similarity, but for exact terminology, codes, table fields, etc., keyword retrieval (e.g., BM25) is often more robust. Therefore, Hybrid Search is widely adopted in engineering practice:Pure vector retrieval excels at capturing semantic similarity, but for exact terminology, codes, table fields, etc., keyword retrieval (e.g., BM25) is often more robust. Therefore, Hybrid Search is widely adopted in engineering practice:
Initial retrieval results often include many "marginally relevant" or "redundant" document chunks, requiring re‑ranking to improve the quality of the final Top‑K:Initial retrieval results often include many "marginally relevant" or "redundant" document chunks, requiring re‑ranking to improve the quality of the final Top‑K:
In more advanced practice, retrieval and generation are no longer a one‑way pipeline but form a closed loop:In more advanced practice, retrieval and generation are no longer a one‑way pipeline but form a closed loop:
The final link is the generation layer, which directly determines the user experience. The goal here is not to let the model "freely improvise," but to have it produce clear, bounded, and well‑cited answers under the constraints of the retrieved evidence.The final link is the generation layer, which directly determines the user experience. The goal here is not to let the model "freely improvise," but to have it produce clear, bounded, and well‑cited answers under the constraints of the retrieved evidence.
In a RAG architecture, the LLM receives not only the user's question but also multiple retrieved document chunks and system instructions. The system typically:In a RAG architecture, the LLM receives not only the user's question but also multiple retrieved document chunks and system instructions. The system typically:
To facilitate auditing and traceability — especially in high‑risk domains like legal, medical, financial, and internal corporate policy — answers often need to include explicit citations:To facilitate auditing and traceability — especially in high‑risk domains like legal, medical, financial, and internal corporate policy — answers often need to include explicit citations:
To further improve results in challenging scenarios, more complex RAG variants are used in practice:To further improve results in challenging scenarios, more complex RAG variants are used in practice:
If RAG primarily addresses "how to look up information in large‑scale unstructured documents," then the structured data and knowledge graph layer is more about "how to elegantly leverage structured knowledge in databases, reporting systems, and graph databases."If RAG primarily addresses "how to look up information in large‑scale unstructured documents," then the structured data and knowledge graph layer is more about "how to elegantly leverage structured knowledge in databases, reporting systems, and graph databases."
In enterprise environments, the truly critical business data — orders, customers, contracts, inventory, behavioral logs — often resides in relational databases, data warehouses, OLAP engines, or graph databases. These systems are already very mature in terms of query capability, computational efficiency, and auditing, but for business users, writing SQL or DSL directly still has a fairly high barrier to entry. Text‑to‑SQL / Text‑to‑DSL and knowledge graph Q&A and reasoning aim to let the LLM serve as a "natural language interface" and "reasoning collaboration partner" without disrupting the stability of these systems.In enterprise environments, the truly critical business data — orders, customers, contracts, inventory, behavioral logs — often resides in relational databases, data warehouses, OLAP engines, or graph databases. These systems are already very mature in terms of query capability, computational efficiency, and auditing, but for business users, writing SQL or DSL directly still has a fairly high barrier to entry. Text‑to‑SQL / Text‑to‑DSL and knowledge graph Q&A and reasoning aim to let the LLM serve as a "natural language interface" and "reasoning collaboration partner" without disrupting the stability of these systems.
The core of this layer is transforming the LLM from "someone who directly gives answers" into "an assistant that can query databases and graph databases":The core of this layer is transforming the LLM from "someone who directly gives answers" into "an assistant that can query databases and graph databases":
Typical solutions are usually "LLM + specialized components" multi‑module architectures:Typical solutions are usually "LLM + specialized components" multi‑module architectures:
The goal of database Q&A is to let business users "ask data questions in natural language" while the system automatically handles query generation, execution, and explanation behind the scenes. Doing this well requires balancing semantic accuracy, syntactic correctness, and execution safety.The goal of database Q&A is to let business users "ask data questions in natural language" while the system automatically handles query generation, execution, and explanation behind the scenes. Doing this well requires balancing semantic accuracy, syntactic correctness, and execution safety.
In the most basic pipeline, the system needs to:In the most basic pipeline, the system needs to:
After query execution, the system must also turn "a cold result set" into "understandable insights":After query execution, the system must also turn "a cold result set" into "understandable insights":
Because LLM‑generated SQL is highly flexible, a layer of security and governance mechanisms is essential:Because LLM‑generated SQL is highly flexible, a layer of security and governance mechanisms is essential:
Knowledge graphs attempt to organize knowledge scattered across text, tables, and logs into a structured network of "entities–relationships–attributes–events," thereby better supporting relationship exploration, multi‑hop reasoning, and complex Q&A. In this direction, LLMs complement traditional information extraction and graph databases well.Knowledge graphs attempt to organize knowledge scattered across text, tables, and logs into a structured network of "entities–relationships–attributes–events," thereby better supporting relationship exploration, multi‑hop reasoning, and complex Q&A. In this direction, LLMs complement traditional information extraction and graph databases well.
Constructing a knowledge graph typically uses a multi‑stage pipeline:Constructing a knowledge graph typically uses a multi‑stage pipeline:
Once the graph is built, the graph database handles efficient storage and retrieval, while the LLM can play the role of "natural language entry point + reasoning controller":Once the graph is built, the graph database handles efficient storage and retrieval, while the LLM can play the role of "natural language entry point + reasoning controller":
In larger‑scale enterprise or industry‑level applications, knowledge graphs often serve as a "domain knowledge platform":In larger‑scale enterprise or industry‑level applications, knowledge graphs often serve as a "domain knowledge platform":
The shared goal of this layer is to upgrade "the model can talk" to "the model can both talk and truly connect to the enterprise's real data and knowledge assets." When RAG, Text‑to‑SQL, knowledge graphs, and traditional data infrastructure are effectively integrated, AI systems can maintain both intelligence and flexibility in complex business environments while also possessing controllability, explainability, and long‑term evolution capability.The shared goal of this layer is to upgrade "the model can talk" to "the model can both talk and truly connect to the enterprise's real data and knowledge assets." When RAG, Text‑to‑SQL, knowledge graphs, and traditional data infrastructure are effectively integrated, AI systems can maintain both intelligence and flexibility in complex business environments while also possessing controllability, explainability, and long‑term evolution capability.
In previous chapters, we focused more on "what the model can do": interpreting images, writing code, conversing with users. But in real-world large model systems, mere "capability" is far from enough: how do we prove these capabilities are stable, reliable, and controllable? How do we ensure outputs align with values and compliance requirements? How do we continuously monitor, iterate, and regress during long-term operations?In previous chapters, we focused more on "what the model can do": interpreting images, writing code, conversing with users. But in real-world large model systems, mere "capability" is far from enough: how do we prove these capabilities are stable, reliable, and controllable? How do we ensure outputs align with values and compliance requirements? How do we continuously monitor, iterate, and regress during long-term operations?
This layer is concerned with: capability evaluation & benchmarks, value alignment & training, content safety & compliance, and robustness & hallucination control — together forming a sustainable "infrastructure layer" for large model operations.This layer is concerned with: capability evaluation & benchmarks, value alignment & training, content safety & compliance, and robustness & hallucination control — together forming a sustainable "infrastructure layer" for large model operations.
From a product perspective, these capabilities span the entire model lifecycle: models require standard benchmarks and professional evaluation during the lab phase; they must pass alignment training and safety reviews before launch; post-launch they rely on content safety gateways, log auditing, and A/B testing for continuous monitoring; and when facing new scenarios and threats, they return to the evaluation and alignment stages for retraining and re-validation. Below, we explore four dimensions: capability evaluation & benchmarks, value alignment & training, content safety & compliance, and robustness & hallucination control.From a product perspective, these capabilities span the entire model lifecycle: models require standard benchmarks and professional evaluation during the lab phase; they must pass alignment training and safety reviews before launch; post-launch they rely on content safety gateways, log auditing, and A/B testing for continuous monitoring; and when facing new scenarios and threats, they return to the evaluation and alignment stages for retraining and re-validation. Below, we explore four dimensions: capability evaluation & benchmarks, value alignment & training, content safety & compliance, and robustness & hallucination control.
In the process of developing and deploying large models, capability evaluation & benchmarks is the critical link that transforms "model capability" into "observable signals": it must answer both "how good is this model overall" and "how does it perform in a specific domain or real business scenario." On one hand, we use standardized benchmark suites and automated evaluation systems to measure model performance on general dimensions such as language understanding & generation, reasoning & math, and knowledge & factuality. On the other hand, we need to build specialized evaluations for domains like healthcare, law, finance, and education, and continuously validate and refine them through real user conversations, A/B testing, and business metrics (Task Success Rate, CSAT, ticket closure rate, etc.). Overall, this layer ultimately crystallizes into an internal capability evaluation platform and an external-facing "capability specification," providing a unified decision-making basis for model selection across multiple versions, tenants, and scenarios. Below, we elaborate from three perspectives: scenarios, principles, and models.In the process of developing and deploying large models, capability evaluation & benchmarks is the critical link that transforms "model capability" into "observable signals": it must answer both "how good is this model overall" and "how does it perform in a specific domain or real business scenario." On one hand, we use standardized benchmark suites and automated evaluation systems to measure model performance on general dimensions such as language understanding & generation, reasoning & math, and knowledge & factuality. On the other hand, we need to build specialized evaluations for domains like healthcare, law, finance, and education, and continuously validate and refine them through real user conversations, A/B testing, and business metrics (Task Success Rate, CSAT, ticket closure rate, etc.). Overall, this layer ultimately crystallizes into an internal capability evaluation platform and an external-facing "capability specification," providing a unified decision-making basis for model selection across multiple versions, tenants, and scenarios. Below, we elaborate from three perspectives: scenarios, principles, and models.
The capability evaluation system can be viewed as a layered "measurement systems engineering" effort, with core principles including:The capability evaluation system can be viewed as a layered "measurement systems engineering" effort, with core principles including:
These benchmarks emphasize standardization, reproducibility, and comparability, facilitating cross-model and cross-institution horizontal comparisons and external disclosure.These benchmarks emphasize standardization, reproducibility, and comparability, facilitating cross-model and cross-institution horizontal comparisons and external disclosure.
The key to automated evaluation lies in stability and consistency — even if imperfect, as long as "the bias is consistent," it can reliably reflect relative model changes in continuous integration (CI).The key to automated evaluation lies in stability and consistency — even if imperfect, as long as "the bias is consistent," it can reliably reflect relative model changes in continuous integration (CI).
Human evaluation serves both to calibrate automated evaluation and as an important basis for externally "explaining model behavior."Human evaluation serves both to calibrate automated evaluation and as an important basis for externally "explaining model behavior."
In engineering practice, capability evaluation crystallizes into a relatively complete "platform + pipeline + metrics system":In engineering practice, capability evaluation crystallizes into a relatively complete "platform + pipeline + metrics system":
General and domain-specific capability evaluation is the "first layer of the foundation" for the entire evaluation system, with the key focus being: first measuring the model's foundational capabilities on a unified yardstick, then validating its usability and risk in specialized scenarios.General and domain-specific capability evaluation is the "first layer of the foundation" for the entire evaluation system, with the key focus being: first measuring the model's foundational capabilities on a unified yardstick, then validating its usability and risk in specialized scenarios.
In general capability evaluation, tasks are typically broken down into three dimensions: language understanding & generation, reasoning & math, and knowledge & factuality. The first examines whether the model can accurately understand context, control style, and output coherent text through reading comprehension, summarization, translation, and dialogue quality tasks; the second evaluates the model's ability in complex reasoning chains and program structure through arithmetic, multi-step reasoning, and code/logic problems; the third measures knowledge coverage and factuality through fact-based QA and open-domain QA. In domain-specific evaluation, industry experts need to be involved in data design: for example, in medical Q&A, setting contexts such as medical history and lab results, requiring the model to provide risk warnings and boundaries for medical advice in its responses; in legal tasks, designing statute retrieval, case comparison, and legal applicability analysis; in finance and education, focusing on compliance disclosures and instructional guidance. This layer of evaluation often combines standard benchmark suites with custom-built datasets, pursuing both comparability and business relevance.In general capability evaluation, tasks are typically broken down into three dimensions: language understanding & generation, reasoning & math, and knowledge & factuality. The first examines whether the model can accurately understand context, control style, and output coherent text through reading comprehension, summarization, translation, and dialogue quality tasks; the second evaluates the model's ability in complex reasoning chains and program structure through arithmetic, multi-step reasoning, and code/logic problems; the third measures knowledge coverage and factuality through fact-based QA and open-domain QA. In domain-specific evaluation, industry experts need to be involved in data design: for example, in medical Q&A, setting contexts such as medical history and lab results, requiring the model to provide risk warnings and boundaries for medical advice in its responses; in legal tasks, designing statute retrieval, case comparison, and legal applicability analysis; in finance and education, focusing on compliance disclosures and instructional guidance. This layer of evaluation often combines standard benchmark suites with custom-built datasets, pursuing both comparability and business relevance.
When the scale of tasks and the number of model versions grow rapidly, relying solely on human evaluation can no longer meet evaluation needs — at this point, an automated evaluation system is required to achieve scalability and high-frequency regression.When the scale of tasks and the number of model versions grow rapidly, relying solely on human evaluation can no longer meet evaluation needs — at this point, an automated evaluation system is required to achieve scalability and high-frequency regression.
One approach uses traditional rule-based metrics: for translation and summarization tasks, comparing against reference answers using BLEU / ROUGE / BERTScore; for coding tasks, using Pass@k to test whether at least one of multiple generated samples passes unit tests. These metrics are simple to implement and highly automatable, but are insensitive to answer diversity and stylistic nuances. Another more representative approach is LLM-as-a-Judge: using a stronger or specially trained model as a "scoring referee," performing dimensional scoring or pairwise ranking of the evaluated model's outputs based on predefined scoring rubrics. This enables efficient automated evaluation even in open-ended QA and dialogue tasks where there are no standard answers and responses are diverse. In practical engineering, the scoring criteria and prompts for LLM-as-a-Judge need to be calibrated and iterated against human-annotated data to ensure consistency with human judges.One approach uses traditional rule-based metrics: for translation and summarization tasks, comparing against reference answers using BLEU / ROUGE / BERTScore; for coding tasks, using Pass@k to test whether at least one of multiple generated samples passes unit tests. These metrics are simple to implement and highly automatable, but are insensitive to answer diversity and stylistic nuances. Another more representative approach is LLM-as-a-Judge: using a stronger or specially trained model as a "scoring referee," performing dimensional scoring or pairwise ranking of the evaluated model's outputs based on predefined scoring rubrics. This enables efficient automated evaluation even in open-ended QA and dialogue tasks where there are no standard answers and responses are diverse. In practical engineering, the scoring criteria and prompts for LLM-as-a-Judge need to be calibrated and iterated against human-annotated data to ensure consistency with human judges.
No matter how comprehensive offline metrics are, they can only approximate real user experience. To close the capability evaluation loop to the business, two types of methods are needed: human evaluation and online experiments.No matter how comprehensive offline metrics are, they can only approximate real user experience. To close the capability evaluation loop to the business, two types of methods are needed: human evaluation and online experiments.
On the human evaluation side, the most common approach is Pairwise Comparison: annotators, without knowing the model identity, make preference selections or score A/B responses based on dimensions such as helpful / honest / harmless, producing high-quality preference data. This data is used both for direct evaluation and for training reward models in RLHF / RLAIF. On the business side, online A/B testing compares the impact of different models, prompts, and policy configuration versions on key metrics such as task completion rate, customer satisfaction (CSAT), and ticket closure rate, supplemented by user conversation log replay and manual spot checks, to continuously monitor the model's real-world performance post-launch. The output of this layer of evaluation in turn feeds back to guide the focus and weight adjustments of the capability evaluation platform, forming a closed loop of "offline metrics — human evaluation — online metrics."On the human evaluation side, the most common approach is Pairwise Comparison: annotators, without knowing the model identity, make preference selections or score A/B responses based on dimensions such as helpful / honest / harmless, producing high-quality preference data. This data is used both for direct evaluation and for training reward models in RLHF / RLAIF. On the business side, online A/B testing compares the impact of different models, prompts, and policy configuration versions on key metrics such as task completion rate, customer satisfaction (CSAT), and ticket closure rate, supplemented by user conversation log replay and manual spot checks, to continuously monitor the model's real-world performance post-launch. The output of this layer of evaluation in turn feeds back to guide the focus and weight adjustments of the capability evaluation platform, forming a closed loop of "offline metrics — human evaluation — online metrics."
After acquiring powerful foundational capabilities, for a large model to become a "safe, reliable, and controllable" product, it must also undergo value alignment & training. This layer is no longer concerned with whether the model "can answer," but with "whether the answer is helpful, honest, and harmless" and "how it should speak in different roles and industries." From an engineering perspective, the alignment process roughly involves three steps: first, clearly defining alignment goal definitions (What to Align) through documentation and specifications, decomposing Helpful, Honest, and Harmless into annotatable and trainable standards; second, constructing broad-coverage instruction data and safety data, covering normal tasks, gray-area cases, and inappropriate responses; and finally, through methods such as SFT, RLHF / RLAIF, and refusal/redirection policy modeling, embedding these preferences and rules into model behavior, supplemented by upstream dialogue management and policy engines to achieve end-to-end safety alignment. Below, we again elaborate from three perspectives: scenarios, principles, and models.After acquiring powerful foundational capabilities, for a large model to become a "safe, reliable, and controllable" product, it must also undergo value alignment & training. This layer is no longer concerned with whether the model "can answer," but with "whether the answer is helpful, honest, and harmless" and "how it should speak in different roles and industries." From an engineering perspective, the alignment process roughly involves three steps: first, clearly defining alignment goal definitions (What to Align) through documentation and specifications, decomposing Helpful, Honest, and Harmless into annotatable and trainable standards; second, constructing broad-coverage instruction data and safety data, covering normal tasks, gray-area cases, and inappropriate responses; and finally, through methods such as SFT, RLHF / RLAIF, and refusal/redirection policy modeling, embedding these preferences and rules into model behavior, supplemented by upstream dialogue management and policy engines to achieve end-to-end safety alignment. Below, we again elaborate from three perspectives: scenarios, principles, and models.
Value alignment can be understood as "constraining the model's behavioral space with human and organizational values," with core principles including:Value alignment can be understood as "constraining the model's behavioral space with human and organizational values," with core principles including:
These goals are written into annotation guidelines and policy documents, serving as unified standards for subsequent data construction, reward modeling, and evaluation.These goals are written into annotation guidelines and policy documents, serving as unified standards for subsequent data construction, reward modeling, and evaluation.
In system design, value alignment typically manifests as a combination of "bottom-layer alignment training + upper-layer policy guardrails":In system design, value alignment typically manifests as a combination of "bottom-layer alignment training + upper-layer policy guardrails":
The first step in value alignment is translating "abstract values" into signals that the model can learn — and this depends on alignment goal definition and training data construction.The first step in value alignment is translating "abstract values" into signals that the model can learn — and this depends on alignment goal definition and training data construction.
At the alignment goal level, teams typically produce a detailed set of behavioral specification documents, decomposing Helpful / Honest / Harmless into specific provisions, such as: prohibiting specific step-by-step instructions for certain high-risk operations, requiring disclaimers and risk warnings for medical/legal advice, maintaining neutrality and multi-perspective presentation on controversial topics, etc. Then, in the instruction data phase, diverse tasks and ideal responses are built around these indicators, covering scenarios like chatting, writing, coding, and Q&A, incorporating multilingual and multicultural contexts. In the safety data phase, paired "good/bad response" examples are constructed targeting harmful content, high-risk domains, and gray zones, providing training material for subsequent preference learning and safety classifiers. Through this approach, value objectives are "translated" into actual data distributions, becoming signals that model training can directly perceive.At the alignment goal level, teams typically produce a detailed set of behavioral specification documents, decomposing Helpful / Honest / Harmless into specific provisions, such as: prohibiting specific step-by-step instructions for certain high-risk operations, requiring disclaimers and risk warnings for medical/legal advice, maintaining neutrality and multi-perspective presentation on controversial topics, etc. Then, in the instruction data phase, diverse tasks and ideal responses are built around these indicators, covering scenarios like chatting, writing, coding, and Q&A, incorporating multilingual and multicultural contexts. In the safety data phase, paired "good/bad response" examples are constructed targeting harmful content, high-risk domains, and gray zones, providing training material for subsequent preference learning and safety classifiers. Through this approach, value objectives are "translated" into actual data distributions, becoming signals that model training can directly perceive.
With alignment goals and data in place, the next step is to write these goals into model behavior through a multi-stage training process.With alignment goals and data in place, the next step is to write these goals into model behavior through a multi-stage training process.
In the SFT stage, the model undergoes supervised fine-tuning on high-quality human demonstration data — akin to "textbook-style learning": it determines the model's tone, structure, and standard problem-solving paradigms for the vast majority of normal requests. Subsequently, RLHF / RLAIF is used for preference optimization: first, a reward model is trained using preference labels produced by human annotation or a larger LLM, then policy optimization algorithms (such as PPO) are used to adjust the model so that it tends to achieve higher rewards during generation. In this way, the model not only "knows what the correct answer looks like" but also knows "which type of answer better aligns with human preferences and safety requirements." On top of this, various refusal & redirection policies are specifically modeled: for questions that are clearly illegal, extremely high-risk, or unsuitable for AI to answer, the model should learn to give clear refusals with explanations and provide safe alternative paths (such as helplines, professional consultation, etc.), rather than simply remaining silent or casually deflecting.In the SFT stage, the model undergoes supervised fine-tuning on high-quality human demonstration data — akin to "textbook-style learning": it determines the model's tone, structure, and standard problem-solving paradigms for the vast majority of normal requests. Subsequently, RLHF / RLAIF is used for preference optimization: first, a reward model is trained using preference labels produced by human annotation or a larger LLM, then policy optimization algorithms (such as PPO) are used to adjust the model so that it tends to achieve higher rewards during generation. In this way, the model not only "knows what the correct answer looks like" but also knows "which type of answer better aligns with human preferences and safety requirements." On top of this, various refusal & redirection policies are specifically modeled: for questions that are clearly illegal, extremely high-risk, or unsuitable for AI to answer, the model should learn to give clear refusals with explanations and provide safe alternative paths (such as helplines, professional consultation, etc.), rather than simply remaining silent or casually deflecting.
Even when the underlying model has undergone sufficient alignment training, a policy layer and alignment platform are still needed in real-world systems to achieve finer-grained controllability and evolvability.Even when the underlying model has undergone sufficient alignment training, a policy layer and alignment platform are still needed in real-world systems to achieve finer-grained controllability and evolvability.
The policy layer typically includes intent recognition, risk assessment, and routing logic: when user input reaches the system, a lightweight model first determines its intent, domain, and risk level, then decides whether to directly invoke the large model, whether additional safety filtering is needed, or whether it should fall into a templated response or human handoff channel. For different industries and customers, the policy layer can load different policy configurations, enabling customization of sensitive categories, refusal styles, and brand tone. Meanwhile, the internal alignment platform manages all alignment-related assets: annotation/scoring tools, reward model versions, policy change records, online A/B results, etc., allowing teams to rapidly iterate and canary-release alignment policies without frequently retraining the base model, thereby maintaining continuous control over model behavior.The policy layer typically includes intent recognition, risk assessment, and routing logic: when user input reaches the system, a lightweight model first determines its intent, domain, and risk level, then decides whether to directly invoke the large model, whether additional safety filtering is needed, or whether it should fall into a templated response or human handoff channel. For different industries and customers, the policy layer can load different policy configurations, enabling customization of sensitive categories, refusal styles, and brand tone. Meanwhile, the internal alignment platform manages all alignment-related assets: annotation/scoring tools, reward model versions, policy change records, online A/B results, etc., allowing teams to rapidly iterate and canary-release alignment policies without frequently retraining the base model, thereby maintaining continuous control over model behavior.
As large models become embedded in search, dialogue, content creation, social platforms, and even enterprise internal systems, content safety & compliance has shifted from an "add-on feature" to a "barrier to entry." This layer is concerned with: whether the model generates illegal or harmful content when producing text, images, audio, and video; whether the system complies with the laws and regulations of the countries/regions and industries in which it operates when processing user data; and whether it can provide clear, traceable evidence chains when facing audits and regulatory oversight. To this end, we need to build a complete technical and governance system covering multimodal content moderation, regional & industry compliance, and local privacy & data protection, and package it into product forms such as SaaS content safety services, enterprise compliance middle platforms, and industry security gateways. Below, we again elaborate from three perspectives: scenarios, principles, and models.As large models become embedded in search, dialogue, content creation, social platforms, and even enterprise internal systems, content safety & compliance has shifted from an "add-on feature" to a "barrier to entry." This layer is concerned with: whether the model generates illegal or harmful content when producing text, images, audio, and video; whether the system complies with the laws and regulations of the countries/regions and industries in which it operates when processing user data; and whether it can provide clear, traceable evidence chains when facing audits and regulatory oversight. To this end, we need to build a complete technical and governance system covering multimodal content moderation, regional & industry compliance, and local privacy & data protection, and package it into product forms such as SaaS content safety services, enterprise compliance middle platforms, and industry security gateways. Below, we again elaborate from three perspectives: scenarios, principles, and models.
The underlying principles of content safety & compliance can be divided into three layers: policy, filtering, and privacy:The underlying principles of content safety & compliance can be divided into three layers: policy, filtering, and privacy:
From a product and system design perspective, content safety & compliance ultimately evolves into a series of reusable "safety services and middle platforms":From a product and system design perspective, content safety & compliance ultimately evolves into a series of reusable "safety services and middle platforms":
A real content safety system must first be able to "understand" content from different channels and modalities, and then enforce policies on every request and response.A real content safety system must first be able to "understand" content from different channels and modalities, and then enforce policies on every request and response.
In multimodal moderation, the system typically builds multiple detection models for text, images, video, etc.: text-side models identify sensitive keywords, contextual semantics, and veiled expressions; image and video-side models detect violence, pornography, minors, hate symbols, illegal items, and other content, combining OCR, ASR, and visual features for joint judgment where necessary. The policy engine binds these model outputs to regulatory requirements: for example, if a region has stricter restrictions on gambling or political content, the sensitivity of related detection categories can be raised in the corresponding policy template, or content matching these categories can be forced into human review. By transforming abstract rules into rule chains, thresholds, and actions (allow/block/human review/mask), the Policy Engine makes compliance requirements truly "operational."In multimodal moderation, the system typically builds multiple detection models for text, images, video, etc.: text-side models identify sensitive keywords, contextual semantics, and veiled expressions; image and video-side models detect violence, pornography, minors, hate symbols, illegal items, and other content, combining OCR, ASR, and visual features for joint judgment where necessary. The policy engine binds these model outputs to regulatory requirements: for example, if a region has stricter restrictions on gambling or political content, the sensitivity of related detection categories can be raised in the corresponding policy template, or content matching these categories can be forced into human review. By transforming abstract rules into rule chains, thresholds, and actions (allow/block/human review/mask), the Policy Engine makes compliance requirements truly "operational."
Single-point interception can hardly cover all risks, so content safety systems generally adopt a pre-event — in-event — post-event three-layer defense design.Single-point interception can hardly cover all risks, so content safety systems generally adopt a pre-event — in-event — post-event three-layer defense design.
In the pre-event stage, the system performs rapid detection on user input, directly rejecting or rewriting clearly violating or highly sensitive prompts, guiding users to ask questions in a safe manner; for borderline attempts and ambiguous requests, it can also proactively add disclaimers and risk warnings. In the in-event stage, model outputs pass through a real-time safety filtering component: this component uses text classification and rule matching to clip, replace, or trigger refusal flows for potentially high-risk outputs, ensuring that the content ultimately presented to the user falls within an acceptable range. In the post-event stage, through log auditing and sampling mechanisms, safety teams or trusted automated systems periodically replay and review conversations, analyze false positives, false negatives, and emerging risk patterns, and accordingly update policies, training data, and detection models. This forms a continuously evolving safety closed loop rather than a "one-time configuration."In the pre-event stage, the system performs rapid detection on user input, directly rejecting or rewriting clearly violating or highly sensitive prompts, guiding users to ask questions in a safe manner; for borderline attempts and ambiguous requests, it can also proactively add disclaimers and risk warnings. In the in-event stage, model outputs pass through a real-time safety filtering component: this component uses text classification and rule matching to clip, replace, or trigger refusal flows for potentially high-risk outputs, ensuring that the content ultimately presented to the user falls within an acceptable range. In the post-event stage, through log auditing and sampling mechanisms, safety teams or trusted automated systems periodically replay and review conversations, analyze false positives, false negatives, and emerging risk patterns, and accordingly update policies, training data, and detection models. This forms a continuously evolving safety closed loop rather than a "one-time configuration."
In highly sensitive industries, "not outputting harmful content" is far from enough — one must also prove that "internal use of user data is equally safe, compliant, and traceable."In highly sensitive industries, "not outputting harmful content" is far from enough — one must also prove that "internal use of user data is equally safe, compliant, and traceable."
Privacy protection begins the moment data enters the system: anonymization and desensitization are performed as early as the collection and storage stages, ensuring that even if logs are leaked, they are difficult to directly associate with specific individuals. During training, differential privacy, sampling strategies, or federated learning reduce the impact and leakage risk of individual user data on the final model. For model inference traffic, a security gateway provides unified access control: all requests and responses pass through the gateway's content inspection, permission verification, and audit logging, with different access policies and data views applied as needed based on business line and user role. Ultimately, these logs and policy change records crystallize into an "evidence chain" viewable by internal auditors and external regulators, enabling the enterprise to not only be compliant in fact but also "provably compliant" in form.Privacy protection begins the moment data enters the system: anonymization and desensitization are performed as early as the collection and storage stages, ensuring that even if logs are leaked, they are difficult to directly associate with specific individuals. During training, differential privacy, sampling strategies, or federated learning reduce the impact and leakage risk of individual user data on the final model. For model inference traffic, a security gateway provides unified access control: all requests and responses pass through the gateway's content inspection, permission verification, and audit logging, with different access policies and data views applied as needed based on business line and user role. Ultimately, these logs and policy change records crystallize into an "evidence chain" viewable by internal auditors and external regulators, enabling the enterprise to not only be compliant in fact but also "provably compliant" in form.
When deep learning and large models move from "recommendation ads and natural language understanding" toward scientific problems themselves, the goal is no longer just predicting a metric or performing classification, but genuinely participating in discovering patterns, designing experiments, and accelerating simulation and reasoning. AI4Science seeks to combine "statistical pattern recognition" with "physical laws / biochemical principles / mathematical structures," enabling models to serve as "programmable scientific assistants" in molecular design, protein engineering, materials discovery, physics simulation, mathematical reasoning, and beyond.When deep learning and large models move from "recommendation ads and natural language understanding" toward scientific problems themselves, the goal is no longer just predicting a metric or performing classification, but genuinely participating in discovering patterns, designing experiments, and accelerating simulation and reasoning. AI4Science seeks to combine "statistical pattern recognition" with "physical laws / biochemical principles / mathematical structures," enabling models to serve as "programmable scientific assistants" in molecular design, protein engineering, materials discovery, physics simulation, mathematical reasoning, and beyond.
In engineering practice, this layer connects on one end to "traditional scientific infrastructure" such as quantum chemistry software, molecular dynamics (MD), CFD/FEA simulators, automated theorem provers, literature databases, and robotic labs, and on the other end to the real-world R&D workflows of pharmaceutical companies, materials enterprises, energy firms, and research institutions. The following discussion unfolds from three perspectives — scenarios, principles, and models — and further subdivides several key directions.In engineering practice, this layer connects on one end to "traditional scientific infrastructure" such as quantum chemistry software, molecular dynamics (MD), CFD/FEA simulators, automated theorem provers, literature databases, and robotic labs, and on the other end to the real-world R&D workflows of pharmaceutical companies, materials enterprises, energy firms, and research institutions. The following discussion unfolds from three perspectives — scenarios, principles, and models — and further subdivides several key directions.
Starting from this layer, traditional scientific computing becomes deeply intertwined with deep learning and large models: both respecting the strict constraints of physics / chemistry / biology / mathematics, and leveraging the powerful data-driven fitting capabilities to improve efficiency. The ultimate goal is for AI to become a "collaborator" in scientific research, not merely a predictive black box.Starting from this layer, traditional scientific computing becomes deeply intertwined with deep learning and large models: both respecting the strict constraints of physics / chemistry / biology / mathematics, and leveraging the powerful data-driven fitting capabilities to improve efficiency. The ultimate goal is for AI to become a "collaborator" in scientific research, not merely a predictive black box.
------
In traditional drug development, going from target discovery to clinical trials often takes 10+ years and billions of dollars, with a huge portion of that time and cost consumed in the early stages of molecular design, property prediction, and virtual screening. AI-driven molecular modeling and drug discovery aims to accelerate this process using data-driven + generative modeling: starting from structures or textual descriptions, predict molecular properties and ADMET, design candidate compounds targeting specific targets, and significantly reduce the burden of wet-lab experiments through multi-objective optimization and virtual screening.In traditional drug development, going from target discovery to clinical trials often takes 10+ years and billions of dollars, with a huge portion of that time and cost consumed in the early stages of molecular design, property prediction, and virtual screening. AI-driven molecular modeling and drug discovery aims to accelerate this process using data-driven + generative modeling: starting from structures or textual descriptions, predict molecular properties and ADMET, design candidate compounds targeting specific targets, and significantly reduce the burden of wet-lab experiments through multi-objective optimization and virtual screening.
This direction connects on one end to data sources such as quantum chemistry software (DFT, ab initio), bioactivity assays, and HTS (High‑Throughput Screening), and on the other end to internal Small Molecule Design platforms, property prediction SaaS, and materials / chemicals design tools within pharmaceutical companies. The following unfolds from three dimensions — scenarios, principles, and models.This direction connects on one end to data sources such as quantum chemistry software (DFT, ab initio), bioactivity assays, and HTS (High‑Throughput Screening), and on the other end to internal Small Molecule Design platforms, property prediction SaaS, and materials / chemicals design tools within pharmaceutical companies. The following unfolds from three dimensions — scenarios, principles, and models.
Starting from this sub-direction, the drug design workflow is moving from "experts + high-throughput experimentation" toward a closed loop of "experts + models + automated experimentation," where AI not only gives scores but gradually participates in the full cycle from "proposing ideas" to "generating candidates" to "screening and optimization."Starting from this sub-direction, the drug design workflow is moving from "experts + high-throughput experimentation" toward a closed loop of "experts + models + automated experimentation," where AI not only gives scores but gradually participates in the full cycle from "proposing ideas" to "generating candidates" to "screening and optimization."
In drug and materials R&D, a fundamental capability is: given a molecule, quickly and accurately predict its properties and behavior, including quantum chemical properties (energy, orbitals, dipole moment), physicochemical properties (solubility, LogP), and pharmacokinetic / toxicity-related ADMET metrics. The essence of this problem is how to learn, from different forms of molecular representation, a representation that both conforms to chemical principles and possesses generalization ability.In drug and materials R&D, a fundamental capability is: given a molecule, quickly and accurately predict its properties and behavior, including quantum chemical properties (energy, orbitals, dipole moment), physicochemical properties (solubility, LogP), and pharmacokinetic / toxicity-related ADMET metrics. The essence of this problem is how to learn, from different forms of molecular representation, a representation that both conforms to chemical principles and possesses generalization ability.
Typical model pathways: use DimeNet / SchNet / PhysNet / GNNs to extract high-dimensional representations from molecular structures, then simultaneously predict multiple properties via multi-task learning; conduct pretraining on large-scale public or internal corporate data to improve modeling capability in small-data scenarios. Externally, these are offered as ADMET prediction SaaS or internal platform APIs, providing project teams with rapid "virtual experiment" capabilities.Typical model pathways: use DimeNet / SchNet / PhysNet / GNNs to extract high-dimensional representations from molecular structures, then simultaneously predict multiple properties via multi-task learning; conduct pretraining on large-scale public or internal corporate data to improve modeling capability in small-data scenarios. Externally, these are offered as ADMET prediction SaaS or internal platform APIs, providing project teams with rapid "virtual experiment" capabilities.
With reliable molecular representation and property prediction models in place, the next goal is to actively generate "better" molecules: no longer just evaluating given compounds, but directly designing new candidate molecules around targets and property constraints. This direction is commonly referred to as molecular generation and molecular optimization.With reliable molecular representation and property prediction models in place, the next goal is to actively generate "better" molecules: no longer just evaluating given compounds, but directly designing new candidate molecules around targets and property constraints. This direction is commonly referred to as molecular generation and molecular optimization.
In terms of structure generation, research and engineering practice center on three main pathways:In terms of structure generation, research and engineering practice center on three main pathways:
Treat molecules as strings, using VAE, GAN, or autoregressive Transformers to sample new structures in SMILES space; ensure chemical validity through grammatical constraints (e.g., SELFIES) or post-processing.Treat molecules as strings, using VAE, GAN, or autoregressive Transformers to sample new structures in SMILES space; ensure chemical validity through grammatical constraints (e.g., SELFIES) or post-processing.
Models such as GraphVAE, Junction Tree VAE, and GraphAF construct structures directly at the molecular graph or primitive fragment (fragment / motif) level, which is closer to chemical synthesis thinking and facilitates control over rings, groups, and scaffold structures.Models such as GraphVAE, Junction Tree VAE, and GraphAF construct structures directly at the molecular graph or primitive fragment (fragment / motif) level, which is closer to chemical synthesis thinking and facilitates control over rings, groups, and scaffold structures.
Methods such as Diffusion for Molecules perform diffusion and denoising in graph or 3D coordinate space, simultaneously considering spatial conformation, suitable for generating ligands or material units that are sensitive to 3D shape.Methods such as Diffusion for Molecules perform diffusion and denoising in graph or 3D coordinate space, simultaneously considering spatial conformation, suitable for generating ligands or material units that are sensitive to 3D shape.
In terms of molecular optimization, the key is introducing objectives and constraints:In terms of molecular optimization, the key is introducing objectives and constraints:
In terms of productization, these models are often packaged into internal "AI drug design platforms" within pharmaceutical companies: given a target, known lead structures, and optimization directions, the platform automatically proposes batches of candidate molecules; project teams then progressively screen and iterate in combination with experimental, patent, and commercial considerations, forming a "model–experiment–model" closed-loop optimization.In terms of productization, these models are often packaged into internal "AI drug design platforms" within pharmaceutical companies: given a target, known lead structures, and optimization directions, the platform automatically proposes batches of candidate molecules; project teams then progressively screen and iterate in combination with experimental, patent, and commercial considerations, forming a "model–experiment–model" closed-loop optimization.
In life sciences, structure determines function is a near-dogmatic principle: how a protein folds into a three-dimensional structure and how it assembles into complexes with other molecules directly determines its functional behavior in cells. Traditional structural determination relies on experimental techniques such as X‑ray crystallography, NMR, and cryo‑EM, which are time-consuming, expensive, and have massive blind spots where "difficult to crystallize, difficult to resolve" applies. Deep learning models represented by AlphaFold have dramatically advanced the ability to go "directly from sequence to structure," making it possible to obtain high-quality structures at the genome-wide scale.In life sciences, structure determines function is a near-dogmatic principle: how a protein folds into a three-dimensional structure and how it assembles into complexes with other molecules directly determines its functional behavior in cells. Traditional structural determination relies on experimental techniques such as X‑ray crystallography, NMR, and cryo‑EM, which are time-consuming, expensive, and have massive blind spots where "difficult to crystallize, difficult to resolve" applies. Deep learning models represented by AlphaFold have dramatically advanced the ability to go "directly from sequence to structure," making it possible to obtain high-quality structures at the genome-wide scale.
This direction connects on one end to sequence and structure databases such as UniProt / PDB, omics experiments, and structural genomics projects, and on the other end to structure design and analysis platforms in industries such as biopharmaceuticals, synthetic biology, and enzyme engineering. The following unfolds from three perspectives — scenarios, principles, and models — and further breaks down key sub-directions.This direction connects on one end to sequence and structure databases such as UniProt / PDB, omics experiments, and structural genomics projects, and on the other end to structure design and analysis platforms in industries such as biopharmaceuticals, synthetic biology, and enzyme engineering. The following unfolds from three perspectives — scenarios, principles, and models — and further breaks down key sub-directions.
From this sub-direction onward, AI is not only "interpreting" naturally occurring protein structures but also "creating" entirely new protein and complex architectures, moving structural biology from the "passive measurement era" into the "active design era."From this sub-direction onward, AI is not only "interpreting" naturally occurring protein structures but also "creating" entirely new protein and complex architectures, moving structural biology from the "passive measurement era" into the "active design era."
Protein structure prediction is one of the most representative breakthroughs at the intersection of structural biology and AI. The core question is: can we predict 3D structures at near-experimental resolution from sequence alone, with little to no reliance on experimental data? In real-world applications, monomer structures are often just the starting point; more critically, how proteins assemble into complexes with other molecules.Protein structure prediction is one of the most representative breakthroughs at the intersection of structural biology and AI. The core question is: can we predict 3D structures at near-experimental resolution from sequence alone, with little to no reliance on experimental data? In real-world applications, monomer structures are often just the starting point; more critically, how proteins assemble into complexes with other molecules.
In monomer structure prediction, the typical workflow includes:In monomer structure prediction, the typical workflow includes:
In complex and assembly prediction, the problem further extends to "how multiple chains organize and interact in space":In complex and assembly prediction, the problem further extends to "how multiple chains organize and interact in space":
In product practice, structure prediction and assembly are often packaged as cloud services or local toolchains, providing foundational structural information for protein function annotation, interaction network modeling, and drug target validation.In product practice, structure prediction and assembly are often packaged as cloud services or local toolchains, providing foundational structural information for protein function annotation, interaction network modeling, and drug target validation.
After mastering the "sequence → structure" mapping, the next step is the inverse problem: given a structure or functional requirement, how do we design suitable protein sequences and mutation strategies? This is the core of protein design and mutation effect prediction.After mastering the "sequence → structure" mapping, the next step is the inverse problem: given a structure or functional requirement, how do we design suitable protein sequences and mutation strategies? This is the core of protein design and mutation effect prediction.
In protein design, key tasks include:In protein design, key tasks include:
In mutation effect prediction, the focus is on:In mutation effect prediction, the focus is on:
At the engineering and product level, protein design and mutation effect prediction are often integrated as "structural design and optimization modules" within biopharmaceutical / synthetic biology companies: starting from candidate backbone structures, automatically propose multiple rounds of mutation and variant library design plans, forming a data-driven closed loop with high-throughput screening experiments.At the engineering and product level, protein design and mutation effect prediction are often integrated as "structural design and optimization modules" within biopharmaceutical / synthetic biology companies: starting from candidate backbone structures, automatically propose multiple rounds of mutation and variant library design plans, forming a data-driven closed loop with high-throughput screening experiments.
In aerospace, automotive, civil engineering, energy, chemical, and other fields, high-fidelity simulation is a core part of design and verification. However, CFD (Computational Fluid Dynamics), FEA (Finite Element Analysis), molecular dynamics (MD), and various PDE solvers are often computationally expensive, making it difficult to support large-scale parameter sweeps, real-time control, or online optimization. AI-driven physics simulation and surrogate modeling seeks to use deep networks to approximate numerical solvers or operators themselves, achieving orders-of-magnitude acceleration while maintaining physical consistency and interpretability.In aerospace, automotive, civil engineering, energy, chemical, and other fields, high-fidelity simulation is a core part of design and verification. However, CFD (Computational Fluid Dynamics), FEA (Finite Element Analysis), molecular dynamics (MD), and various PDE solvers are often computationally expensive, making it difficult to support large-scale parameter sweeps, real-time control, or online optimization. AI-driven physics simulation and surrogate modeling seeks to use deep networks to approximate numerical solvers or operators themselves, achieving orders-of-magnitude acceleration while maintaining physical consistency and interpretability.
This direction connects on one end to traditional simulation software (ANSYS, Fluent, COMSOL, custom solvers), experimental measurements, and sensor data, and on the other end to engineering design platforms, autonomous driving and aerospace aerodynamic design, and chemical process simulation and optimization systems. The following unfolds from three perspectives — scenarios, principles, and models.This direction connects on one end to traditional simulation software (ANSYS, Fluent, COMSOL, custom solvers), experimental measurements, and sensor data, and on the other end to engineering design platforms, autonomous driving and aerospace aerodynamic design, and chemical process simulation and optimization systems. The following unfolds from three perspectives — scenarios, principles, and models.
Surrogate models and Physics-Informed Neural Networks (PINN) are two complementary paths for AI-enabled physics simulation: the former approximates simulation mappings from data, while the latter constructs learning objectives from physics.Surrogate models and Physics-Informed Neural Networks (PINN) are two complementary paths for AI-enabled physics simulation: the former approximates simulation mappings from data, while the latter constructs learning objectives from physics.
In surrogate model scenarios, the typical workflow is:In surrogate model scenarios, the typical workflow is:
In PINN scenarios, the model no longer relies primarily on large amounts of supervised labels, but instead constructs the loss function by minimizing PDE residuals and boundary condition violations:In PINN scenarios, the model no longer relies primarily on large amounts of supervised labels, but instead constructs the loss function by minimizing PDE residuals and boundary condition violations:
The two can be used in combination: when partial high-fidelity data is available, jointly constrain training with data error + physical residuals to improve accuracy and generalization. In engineering applications, PINN is particularly well-suited for inverse problems and data-driven modeling, such as inferring material parameters, source terms, or defect locations from sensor observations.The two can be used in combination: when partial high-fidelity data is available, jointly constrain training with data error + physical residuals to improve accuracy and generalization. In engineering applications, PINN is particularly well-suited for inverse problems and data-driven modeling, such as inferring material parameters, source terms, or defect locations from sensor observations.
Neural Operators elevate physics modeling from "point-to-point / parameter-to-solution" mappings to the "function-to-function" level: they learn a unified operator approximation for "given a class of PDEs and boundary conditions, solve for the solution field," rather than a specific solution under a single operating condition. This opens new possibilities for generalization across multiple conditions, geometries, and mesh resolutions.Neural Operators elevate physics modeling from "point-to-point / parameter-to-solution" mappings to the "function-to-function" level: they learn a unified operator approximation for "given a class of PDEs and boundary conditions, solve for the solution field," rather than a specific solution under a single operating condition. This opens new possibilities for generalization across multiple conditions, geometries, and mesh resolutions.
In operator learning, typical approaches are:In operator learning, typical approaches are:
In multi-scale modeling scenarios:In multi-scale modeling scenarios:
In engineering practice, Neural Operators are gradually moving from research prototypes to applications, becoming an important technical direction for "accelerated solvers + multi-scale bridging" in CFD, geophysics, climate modeling, and other scenarios.In engineering practice, Neural Operators are gradually moving from research prototypes to applications, becoming an important technical direction for "accelerated solvers + multi-scale bridging" in CFD, geophysics, climate modeling, and other scenarios.
In materials science, a core contradiction is: the design space is nearly infinite, while the cost of experimentation and high-precision computation is extremely high. How to efficiently find candidate materials that meet specific performance requirements within the vast chemical and structural combination space is a key problem in new energy, electronics, structural materials, functional materials, and other fields. AI-driven materials discovery and crystal design, through graph neural networks, generative models, and high-throughput virtual screening, is progressively shifting "trial-and-error" R&D toward "data-driven + inverse design."In materials science, a core contradiction is: the design space is nearly infinite, while the cost of experimentation and high-precision computation is extremely high. How to efficiently find candidate materials that meet specific performance requirements within the vast chemical and structural combination space is a key problem in new energy, electronics, structural materials, functional materials, and other fields. AI-driven materials discovery and crystal design, through graph neural networks, generative models, and high-throughput virtual screening, is progressively shifting "trial-and-error" R&D toward "data-driven + inverse design."
This direction connects on one end to materials databases such as Materials Project, OQMD, AFLOW and DFT / MD computation results, and on the other end to materials R&D platforms in application scenarios such as batteries, photovoltaics, catalysis, semiconductors, and alloys. The following unfolds from three perspectives — scenarios, principles, and models.This direction connects on one end to materials databases such as Materials Project, OQMD, AFLOW and DFT / MD computation results, and on the other end to materials R&D platforms in application scenarios such as batteries, photovoltaics, catalysis, semiconductors, and alloys. The following unfolds from three perspectives — scenarios, principles, and models.
In the materials R&D workflow, fast and reliable property prediction is a foundational capability: given a candidate structure or composition, can we roughly determine whether it is worth further exploration without performing expensive DFT / experimentation? Property prediction models based on GNNs and materials databases make high-throughput virtual screening possible.In the materials R&D workflow, fast and reliable property prediction is a foundational capability: given a candidate structure or composition, can we roughly determine whether it is worth further exploration without performing expensive DFT / experimentation? Property prediction models based on GNNs and materials databases make high-throughput virtual screening possible.
At the property prediction level:At the property prediction level:
In high-throughput virtual screening (HTVS) scenarios, the typical workflow is:In high-throughput virtual screening (HTVS) scenarios, the typical workflow is:
This workflow has entered practical use in multiple domains including battery materials, photovoltaic absorber layers, catalysts, and structural materials, becoming a "front-end screening engine" for materials R&D teams.This workflow has entered practical use in multiple domains including battery materials, photovoltaic absorber layers, catalysts, and structural materials, becoming a "front-end screening engine" for materials R&D teams.
With reliable property prediction and HTVS capabilities in place, the next goal is to directly propose new crystal structures and composition candidates from target properties and constraints — i.e., materials inverse design and generation.With reliable property prediction and HTVS capabilities in place, the next goal is to directly propose new crystal structures and composition candidates from target properties and constraints — i.e., materials inverse design and generation.
In crystal generation, key questions include:In crystal generation, key questions include:
To address these, research and engineering practice commonly adopt:To address these, research and engineering practice commonly adopt:
In inverse design, it is typically combined with surrogate models and optimization methods:In inverse design, it is typically combined with surrogate models and optimization methods:
In engineering applications, inverse design modules are often integrated into materials AI platforms, providing R&D personnel with an interactive interface for "set target properties → system automatically proposes candidate structures," significantly improving the efficiency of new materials exploration.In engineering applications, inverse design modules are often integrated into materials AI platforms, providing R&D personnel with an interactive interface for "set target properties → system automatically proposes candidate structures," significantly improving the efficiency of new materials exploration.
Mathematics is a highly formalized, precisely verifiable language, which gives it both "extremely high difficulty" and "potentially enormous returns" in the AI era. On one hand, complex theorem proving and higher-order reasoning place extremely high demands on model capabilities; on the other hand, the results of mathematical reasoning and symbolic computation can be rigorously verified, making them naturally suited for collaboration with programmatic tools. The goal of AI in mathematics and symbolic reasoning is to build models capable of reliable reasoning and computation within formal systems, and to integrate them into education, research, and engineering applications.Mathematics is a highly formalized, precisely verifiable language, which gives it both "extremely high difficulty" and "potentially enormous returns" in the AI era. On one hand, complex theorem proving and higher-order reasoning place extremely high demands on model capabilities; on the other hand, the results of mathematical reasoning and symbolic computation can be rigorously verified, making them naturally suited for collaboration with programmatic tools. The goal of AI in mathematics and symbolic reasoning is to build models capable of reliable reasoning and computation within formal systems, and to integrate them into education, research, and engineering applications.
This direction connects on one end to interactive theorem provers such as Lean / Coq / Isabelle, computer algebra systems (CAS) such as SymPy / Mathematica / Maple, and large mathematical problem banks and literature corpora; on the other end to mathematics education products, research assistance tools, and formula derivation and risk analysis needs in engineering / finance and other fields. The following unfolds from three perspectives — scenarios, principles, and models.This direction connects on one end to interactive theorem provers such as Lean / Coq / Isabelle, computer algebra systems (CAS) such as SymPy / Mathematica / Maple, and large mathematical problem banks and literature corpora; on the other end to mathematics education products, research assistance tools, and formula derivation and risk analysis needs in engineering / finance and other fields. The following unfolds from three perspectives — scenarios, principles, and models.
Automated Theorem Proving (ATP) and Interactive Theorem Proving (ITP) are important directions at the intersection of mathematics and computer science. The core task for AI in this domain is to automatically construct or assist in constructing proofs within formal systems, reducing the human burden on low-level details and allowing more focus on high-level ideas.Automated Theorem Proving (ATP) and Interactive Theorem Proving (ITP) are important directions at the intersection of mathematics and computer science. The core task for AI in this domain is to automatically construct or assist in constructing proofs within formal systems, reducing the human burden on low-level details and allowing more focus on high-level ideas.
In formal systems:In formal systems:
AI can play multiple roles in this:AI can play multiple roles in this:
AlphaZero‑style provers, GPT‑f, Lean‑Dojo, and other works, by training policy and value networks or language models on large-scale formalized corpora, have achieved automatic proof of a considerable proportion of theorems in systems such as Lean / Coq. In terms of product direction, this capability is expected to evolve into "formal verification assistants" for software / hardware verification, cryptographic protocol analysis, and high-reliability system design.AlphaZero‑style provers, GPT‑f, Lean‑Dojo, and other works, by training policy and value networks or language models on large-scale formalized corpora, have achieved automatic proof of a considerable proportion of theorems in systems such as Lean / Coq. In terms of product direction, this capability is expected to evolve into "formal verification assistants" for software / hardware verification, cryptographic protocol analysis, and high-reliability system design.
Compared to theorem proving, symbolic computation and mathematical problem solving are closer to engineering and education scenarios. The goal is: starting from natural language problems, automatically construct symbolic expressions, perform computations, and produce interpretable solution steps.Compared to theorem proving, symbolic computation and mathematical problem solving are closer to engineering and education scenarios. The goal is: starting from natural language problems, automatically construct symbolic expressions, perform computations, and produce interpretable solution steps.
In this direction, the typical neural–symbolic collaboration workflow is:In this direction, the typical neural–symbolic collaboration workflow is:
This model has several key advantages:This model has several key advantages:
In engineering / finance scenarios, this capability can be extended to the formalization and analysis of complex models: automatically extracting model structures from documents and code, constructing symbolic representations, and performing sensitivity analysis, boundary case analysis, and risk identification.In engineering / finance scenarios, this capability can be extended to the formalization and analysis of complex models: automatically extracting model structures from documents and code, constructing symbolic representations, and performing sensitivity analysis, boundary case analysis, and risk identification.
The preceding sub-directions mostly focus on "single-point capabilities": predicting a property, generating a structure, proving a theorem. However, in real-world scientific and industrial R&D, what is more critical is how to chain these capabilities into complete workflows and connect them with literature, databases, simulation platforms, and automated experimental equipment. The scientific workflow and lab automation direction aims to build integrated Agent + Tools + Robots systems for scientific scenarios, enabling AI to evolve from "knowing how to compute" to "knowing how to run experiments and conduct research."The preceding sub-directions mostly focus on "single-point capabilities": predicting a property, generating a structure, proving a theorem. However, in real-world scientific and industrial R&D, what is more critical is how to chain these capabilities into complete workflows and connect them with literature, databases, simulation platforms, and automated experimental equipment. The scientific workflow and lab automation direction aims to build integrated Agent + Tools + Robots systems for scientific scenarios, enabling AI to evolve from "knowing how to compute" to "knowing how to run experiments and conduct research."
This direction connects on one end to paper and patent databases (e.g., PubMed, arXiv), scientific data warehouses, domain knowledge graphs, and simulation platforms, and on the other end to robotic labs, high-throughput screening equipment, and research process management systems. The following unfolds from three perspectives — scenarios, principles, and models.This direction connects on one end to paper and patent databases (e.g., PubMed, arXiv), scientific data warehouses, domain knowledge graphs, and simulation platforms, and on the other end to robotic labs, high-throughput screening equipment, and research process management systems. The following unfolds from three perspectives — scenarios, principles, and models.
The vast majority of scientific knowledge first appears in the form of papers and reports. For AI to truly participate in research, it must be able to "read papers and extract structured knowledge from them." Scientific literature mining and knowledge base construction is precisely about building queryable, reason-able knowledge infrastructure from unstructured text.The vast majority of scientific knowledge first appears in the form of papers and reports. For AI to truly participate in research, it must be able to "read papers and extract structured knowledge from them." Scientific literature mining and knowledge base construction is precisely about building queryable, reason-able knowledge infrastructure from unstructured text.
In this direction, core tasks include:In this direction, core tasks include:
To achieve these goals, the following are commonly adopted:To achieve these goals, the following are commonly adopted:
The constructed domain knowledge bases and knowledge graphs not only provide smarter retrieval and recommendation services for R&D personnel, but also supply data and prior support for subsequent experimental design and materials / drug inverse design.The constructed domain knowledge bases and knowledge graphs not only provide smarter retrieval and recommendation services for R&D personnel, but also supply data and prior support for subsequent experimental design and materials / drug inverse design.
With literature mining, modeling, and optimization capabilities in place, the next step is to combine these capabilities with automated experimental platforms to build truly Self‑Driving Labs and scientific workflow agents.With literature mining, modeling, and optimization capabilities in place, the next step is to combine these capabilities with automated experimental platforms to build truly Self‑Driving Labs and scientific workflow agents.
In a Self‑Driving Lab, the typical closed-loop workflow is:In a Self‑Driving Lab, the typical closed-loop workflow is:
In the broader scientific workflow agent, this closed loop extends to simulation, data analysis, and report generation:In the broader scientific workflow agent, this closed loop extends to simulation, data analysis, and report generation:
In terms of product form, such systems are typically delivered as platforms: providing a unified interface and API, connecting to literature databases, simulation engines, and experimental equipment, allowing scientists and engineers to set goals at a high level using natural language and visual interfaces, with the remaining steps automatically orchestrated and executed by agents + toolchains.In terms of product form, such systems are typically delivered as platforms: providing a unified interface and API, connecting to literature databases, simulation engines, and experimental equipment, allowing scientists and engineers to set goals at a high level using natural language and visual interfaces, with the remaining steps automatically orchestrated and executed by agents + toolchains.
From this sub-direction onward, AI's role in science truly shifts from "offline analysis tool" to "online research collaborator": not only able to read papers, write code, and compute models, but also to work alongside robots to complete real experiments and discoveries.From this sub-direction onward, AI's role in science truly shifts from "offline analysis tool" to "online research collaborator": not only able to read papers, write code, and compute models, but also to work alongside robots to complete real experiments and discoveries.
Moving large models from the lab to enterprise production is never just about "the model itself being good enough" — it requires a complete set of stable, scalable, and operable platform and engineering systems. This system must span the entire lifecycle: model training and fine-tuning, deployment and inference optimization, data and model operations, monitoring and cost management, security and compliance, as well as platform and application enablement capabilities, threading together otherwise scattered technical components into a sustainable closed loop.Moving large models from the lab to enterprise production is never just about "the model itself being good enough" — it requires a complete set of stable, scalable, and operable platform and engineering systems. This system must span the entire lifecycle: model training and fine-tuning, deployment and inference optimization, data and model operations, monitoring and cost management, security and compliance, as well as platform and application enablement capabilities, threading together otherwise scattered technical components into a sustainable closed loop.
From a business perspective, platform and engineering capabilities often determine whether an organization can use large models "at scale, safely, and at low cost." With the same underlying model, without a sound MLOps system, you may be stuck at the demo and pilot stage; but once a mature platform is in place, enterprises can rapidly replicate and evolve high-quality applications across multiple business units, countries/regions, and industry scenarios. Below, we elaborate across six dimensions: model training and fine-tuning platforms, deployment and inference optimization, data and model operations, monitoring and cost/reliability, security and compliance infrastructure, and upper-layer application and platform capabilities.From a business perspective, platform and engineering capabilities often determine whether an organization can use large models "at scale, safely, and at low cost." With the same underlying model, without a sound MLOps system, you may be stuck at the demo and pilot stage; but once a mature platform is in place, enterprises can rapidly replicate and evolve high-quality applications across multiple business units, countries/regions, and industry scenarios. Below, we elaborate across six dimensions: model training and fine-tuning platforms, deployment and inference optimization, data and model operations, monitoring and cost/reliability, security and compliance infrastructure, and upper-layer application and platform capabilities.
At the base model level, most organizations do not train hundred-billion-parameter models from scratch. Instead, they perform continued pre-training + fine-tuning on open-source or commercial base models. The core question at this layer is: how to efficiently leverage compute and data to "pull" a general-purpose large model closer to a specific industry, enterprise, or task, while ensuring engineering manageability across multiple models and versions.At the base model level, most organizations do not train hundred-billion-parameter models from scratch. Instead, they perform continued pre-training + fine-tuning on open-source or commercial base models. The core question at this layer is: how to efficiently leverage compute and data to "pull" a general-purpose large model closer to a specific industry, enterprise, or task, while ensuring engineering manageability across multiple models and versions.
From an engineering perspective, this layer typically comprises three parts: pre-training and continued pre-training, fine-tuning paradigms and toolchains, and large-scale distributed training infrastructure.From an engineering perspective, this layer typically comprises three parts: pre-training and continued pre-training, fine-tuning paradigms and toolchains, and large-scale distributed training infrastructure.
In terms of product form, this layer often manifests as: base model research and development platforms, enterprise-level "training-as-a-service + customization" services, one-click fine-tuning platforms, and model marketplaces (Model Hub / Model Store), supporting the production path from "general-purpose models" to "thousands of models for thousands of enterprises."In terms of product form, this layer often manifests as: base model research and development platforms, enterprise-level "training-as-a-service + customization" services, one-click fine-tuning platforms, and model marketplaces (Model Hub / Model Store), supporting the production path from "general-purpose models" to "thousands of models for thousands of enterprises."
Pre-training is the "source engineering" of modern large model capabilities: through self-supervised learning on massive unlabeled text, code, and multimodal data, the model gradually acquires language modeling, world knowledge, basic reasoning, and representation learning abilities. On top of this, continued pre-training (especially Domain-adaptive Pretraining, DAPT) takes on the task of "pulling the model toward a specific vertical domain."Pre-training is the "source engineering" of modern large model capabilities: through self-supervised learning on massive unlabeled text, code, and multimodal data, the model gradually acquires language modeling, world knowledge, basic reasoning, and representation learning abilities. On top of this, continued pre-training (especially Domain-adaptive Pretraining, DAPT) takes on the task of "pulling the model toward a specific vertical domain."
In the general-purpose pre-training phase, core considerations include:In the general-purpose pre-training phase, core considerations include:
In the industry continued pre-training (DAPT) phase, the focus shifts to:In the industry continued pre-training (DAPT) phase, the focus shifts to:
In engineering practice, pre-training and continued pre-training run on large-scale distributed frameworks (Megatron-LM, DeepSpeed ZeRO, etc.) and efficient data pipelines (WebDataset / HF Datasets + object storage), forming stable, reusable training pipelines. For cloud vendors or large companies, such pipelines are often encapsulated as internal platforms, supporting periodic incremental pre-training and parallel iteration of multiple industry base models.In engineering practice, pre-training and continued pre-training run on large-scale distributed frameworks (Megatron-LM, DeepSpeed ZeRO, etc.) and efficient data pipelines (WebDataset / HF Datasets + object storage), forming stable, reusable training pipelines. For cloud vendors or large companies, such pipelines are often encapsulated as internal platforms, supporting periodic incremental pre-training and parallel iteration of multiple industry base models.
After having a powerful pre-trained base, the key to making the model "useful for the business" and "behaviorally controllable" lies in the fine-tuning and alignment stages. This includes both traditional supervised fine-tuning (SFT) and instruction fine-tuning, multi-task fine-tuning, and feedback-based reinforcement learning (RLHF / RLAIF).After having a powerful pre-trained base, the key to making the model "useful for the business" and "behaviorally controllable" lies in the fine-tuning and alignment stages. This includes both traditional supervised fine-tuning (SFT) and instruction fine-tuning, multi-task fine-tuning, and feedback-based reinforcement learning (RLHF / RLAIF).
At the fine-tuning paradigm level, it can be roughly divided into:At the fine-tuning paradigm level, it can be roughly divided into:
When the task distribution differs significantly from pre-training, or when ultimate performance is rigidly required and compute is abundant (such as for specific programming language models or specific language / industry dialogue models), directly updating all parameters yields the maximum performance ceiling. However, its cost is high and version management is complex, so it is generally used only for a few core models.When the task distribution differs significantly from pre-training, or when ultimate performance is rigidly required and compute is abundant (such as for specific programming language models or specific language / industry dialogue models), directly updating all parameters yields the maximum performance ceiling. However, its cost is high and version management is complex, so it is generally used only for a few core models.
Through methods such as Adapter, LoRA / QLoRA, Prefix / P-Tuning, only the inserted "small incremental parameter blocks" or low-rank weight deltas are trained, while the original large model weights remain frozen. This brings three engineering advantages:Through methods such as Adapter, LoRA / QLoRA, Prefix / P-Tuning, only the inserted "small incremental parameter blocks" or low-rank weight deltas are trained, while the original large model weights remain frozen. This brings three engineering advantages:
At the behavioral alignment and safety level, RLHF / RLAIF plays a key role:At the behavioral alignment and safety level, RLHF / RLAIF plays a key role:
In terms of toolchains, frameworks such as Hugging Face Transformers + PEFT, TRL / trlx, and DeepSpeed-RLHF have essentially formed a standard industrial workflow from SFT → RM training → RLHF. In product definition, typical implementations at this layer include: model customization / training-as-a-service, one-click fine-tuning platforms, multi-tenant model marketplaces, and industry / enterprise proprietary large model engineering platforms.In terms of toolchains, frameworks such as Hugging Face Transformers + PEFT, TRL / trlx, and DeepSpeed-RLHF have essentially formed a standard industrial workflow from SFT → RM training → RLHF. In product definition, typical implementations at this layer include: model customization / training-as-a-service, one-click fine-tuning platforms, multi-tenant model marketplaces, and industry / enterprise proprietary large model engineering platforms.
After training a large model, how to provide inference services in a highly available, low-latency, scalable, and cost-reducible manner is the second pillar of the AI engineering system. The deployment and inference layer connects GPU / NPU compute clusters on one end and API gateways, enterprise applications, and external-facing platforms on the other. Its core responsibilities include: deployment architecture design, model routing strategies, inference performance optimization, and hardware utilization.After training a large model, how to provide inference services in a highly available, low-latency, scalable, and cost-reducible manner is the second pillar of the AI engineering system. The deployment and inference layer connects GPU / NPU compute clusters on one end and API gateways, enterprise applications, and external-facing platforms on the other. Its core responsibilities include: deployment architecture design, model routing strategies, inference performance optimization, and hardware utilization.
Overall, this layer addresses three questions: what architecture to use for external service, how to make inference faster and cheaper, and how to maintain high availability and governability in multi-model, multi-region, multi-tenant environments.Overall, this layer addresses three questions: what architecture to use for external service, how to make inference faster and cheaper, and how to maintain high availability and governability in multi-model, multi-region, multi-tenant environments.
On the product side, this layer often takes the form of enterprise AI platforms / model service buses, external cloud APIs, unified inference gateways, high-QPS online inference clusters, low-cost batch processing platforms, and compute utilization optimization solutions — the runtime "operating system" that supports the large-scale deployment of large model capabilities.On the product side, this layer often takes the form of enterprise AI platforms / model service buses, external cloud APIs, unified inference gateways, high-QPS online inference clusters, low-cost batch processing platforms, and compute utilization optimization solutions — the runtime "operating system" that supports the large-scale deployment of large model capabilities.
In the early experimentation phase, many teams choose to provide services with a single "large and comprehensive" model as the single entry point: all requests are handled by the same model. This architecture is simple and has low maintenance costs, suitable for POC and low-traffic scenarios. However, as business expands and cost pressure increases, the shortcomings of a single-model architecture quickly become apparent:In the early experimentation phase, many teams choose to provide services with a single "large and comprehensive" model as the single entry point: all requests are handled by the same model. This architecture is simple and has low maintenance costs, suitable for POC and low-traffic scenarios. However, as business expands and cost pressure increases, the shortcomings of a single-model architecture quickly become apparent:
Therefore, mature large model serving architectures tend to evolve into multi-model serving and intelligent routing architectures:Therefore, mature large model serving architectures tend to evolve into multi-model serving and intelligent routing architectures:
Technically, this often employs a combination of Kubernetes + Service Mesh (Istio / Linkerd) + API Gateway (Kong / APISIX / Envoy) + Model Serving Frameworks (vLLM / TGI / Triton / Ray Serve / KServe), forming a service-mesh-based inference platform that supports multi-model, multi-tenant operation as well as traffic governance and canary releases.Technically, this often employs a combination of Kubernetes + Service Mesh (Istio / Linkerd) + API Gateway (Kong / APISIX / Envoy) + Model Serving Frameworks (vLLM / TGI / Triton / Ray Serve / KServe), forming a service-mesh-based inference platform that supports multi-model, multi-tenant operation as well as traffic governance and canary releases.
In large-scale commercial deployment of large models, inference cost is often one of the largest ongoing expenses. How to compress unit request cost (Cost per Request / per Token) and end-to-end latency to an acceptable range while maintaining experience quality is the core technical challenge of the deployment layer.In large-scale commercial deployment of large models, inference cost is often one of the largest ongoing expenses. How to compress unit request cost (Cost per Request / per Token) and end-to-end latency to an acceptable range while maintaining experience quality is the core technical challenge of the deployment layer.
On the model side, common techniques include:On the model side, common techniques include:
Compressing weights and activations from FP16 / BF16 to low-bit formats such as INT8 / INT4 / NF4, significantly reducing GPU memory usage and bandwidth overhead.Compressing weights and activations from FP16 / BF16 to low-bit formats such as INT8 / INT4 / NF4, significantly reducing GPU memory usage and bandwidth overhead.
Removing unimportant weights or channels through structured / unstructured pruning to make the model sparse, and combining with hardware-friendly sparse operators (such as NVIDIA sparse matrix acceleration) to improve inference speed.Removing unimportant weights or channels through structured / unstructured pruning to make the model sparse, and combining with hardware-friendly sparse operators (such as NVIDIA sparse matrix acceleration) to improve inference speed.
Using a large model as a teacher to distill knowledge into a smaller student model or task-specific model, significantly reducing parameter scale while maintaining close task performance — suitable for latency-sensitive online services or edge deployment.Using a large model as a teacher to distill knowledge into a smaller student model or task-specific model, significantly reducing parameter scale while maintaining close task performance — suitable for latency-sensitive online services or edge deployment.
On the system and runtime side, key optimization points include:On the system and runtime side, key optimization points include:
Cache the attention keys and values of historical tokens during autoregressive generation to avoid repeated computation, thus improving efficiency for long conversations and multi-turn requests; combine with chunked computation and dynamic pruning strategies to control GPU memory overhead.Cache the attention keys and values of historical tokens during autoregressive generation to avoid repeated computation, thus improving efficiency for long conversations and multi-turn requests; combine with chunked computation and dynamic pruning strategies to control GPU memory overhead.
Improve overall throughput without significantly increasing P95 latency through dynamic batching, group scheduling, and parallel token generation across multiple requests; combine with streaming output to improve the frontend interaction experience.Improve overall throughput without significantly increasing P95 latency through dynamic batching, group scheduling, and parallel token generation across multiple requests; combine with streaming output to improve the frontend interaction experience.
Use compilers and runtimes (such as TensorRT, TVM, ONNX Runtime, TorchInductor) for operator fusion, memory layout optimization, and static graph compilation, reducing kernel launch and memory access overhead.Use compilers and runtimes (such as TensorRT, TVM, ONNX Runtime, TorchInductor) for operator fusion, memory layout optimization, and static graph compilation, reducing kernel launch and memory access overhead.
Allocate tasks reasonably across heterogeneous resources such as GPU, CPU, NPU, and FPGA based on the computational characteristics and latency requirements of different tasks:Allocate tasks reasonably across heterogeneous resources such as GPU, CPU, NPU, and FPGA based on the computational characteristics and latency requirements of different tasks:
In terms of tools and frameworks, TensorRT-LLM, SgLang, vLLM, FasterTransformer, LMDeploy, DeepSpeed-Inference, and others have formed a relatively mature large model inference acceleration ecosystem. On the business side, these optimizations ultimately manifest as: high-QPS, low-latency online inference clusters, low-cost batch generation platforms, compute utilization optimization solutions, and MaaS / API billing and cost accounting systems.In terms of tools and frameworks, TensorRT-LLM, SgLang, vLLM, FasterTransformer, LMDeploy, DeepSpeed-Inference, and others have formed a relatively mature large model inference acceleration ecosystem. On the business side, these optimizations ultimately manifest as: high-QPS, low-latency online inference clusters, low-cost batch generation platforms, compute utilization optimization solutions, and MaaS / API billing and cost accounting systems.
Once a large model enters production, it is no longer a "one-time delivery" static asset, but rather a dynamic system that requires continuous iteration across five dimensions: data, model, configuration, versioning, and experimentation. The Data / Model Ops layer is the engineering paradigm built around this reality: from data flywheels and model lifecycle management to online experimentation and automated releases, it provides the foundation for sustainable improvement and controlled evolution of model capabilities.Once a large model enters production, it is no longer a "one-time delivery" static asset, but rather a dynamic system that requires continuous iteration across five dimensions: data, model, configuration, versioning, and experimentation. The Data / Model Ops layer is the engineering paradigm built around this reality: from data flywheels and model lifecycle management to online experimentation and automated releases, it provides the foundation for sustainable improvement and controlled evolution of model capabilities.
This layer connects data lakes / warehouses, logging and collection systems on one end, and training platforms, evaluation systems, and online service gateways on the other — it is the hub that closes the "data–model–business feedback" loop.This layer connects data lakes / warehouses, logging and collection systems on one end, and training platforms, evaluation systems, and online service gateways on the other — it is the hub that closes the "data–model–business feedback" loop.
In traditional software development, version upgrades are often driven by development plans; in the large model era, data and feedback become the primary drivers of iteration. The goal of the data flywheel is to turn "model usage → data accumulation → retraining → model upgrade" into an automatically rolling closed loop, so that models get better with use in real business.In traditional software development, version upgrades are often driven by development plans; in the large model era, data and feedback become the primary drivers of iteration. The goal of the data flywheel is to turn "model usage → data accumulation → retraining → model upgrade" into an automatically rolling closed loop, so that models get better with use in real business.
Core components include:Core components include:
In applications such as chatbots, Copilots, search Q&A, and code assistants, every user interaction is a potentially high-value training sample. Through logging systems and event tracking, structurally collect requests, model responses, and user behavior (clicks, adoption or not), and perform privacy desensitization and field trimming at the collection end to avoid introducing additional compliance risks.In applications such as chatbots, Copilots, search Q&A, and code assistants, every user interaction is a potentially high-value training sample. Through logging systems and event tracking, structurally collect requests, model responses, and user behavior (clicks, adoption or not), and perform privacy desensitization and field trimming at the collection end to avoid introducing additional compliance risks.
Filter out the small fraction of samples most valuable for training from massive logs, such as:Filter out the small fraction of samples most valuable for training from massive logs, such as:
Perform manual or semi-automatic annotation on candidate samples (including expected responses, quality ranking, safety labels, etc.), and ensure annotation quality through multiple rounds of quality inspection, review, and spot-checking, providing reliable data for subsequent SFT or RLHF.Perform manual or semi-automatic annotation on candidate samples (including expected responses, quality ranking, safety labels, etc.), and ensure annotation quality through multiple rounds of quality inspection, review, and spot-checking, providing reliable data for subsequent SFT or RLHF.
Periodically add new samples to the training set, perform SFT / DAPT / RLHF and other retraining operations, and simultaneously evaluate both "offline metrics + online performance" through standard evaluation sets and online A/B experiments, ensuring the new version is overall better than the old version and preventing the data flywheel from "veering in the wrong direction."Periodically add new samples to the training set, perform SFT / DAPT / RLHF and other retraining operations, and simultaneously evaluate both "offline metrics + online performance" through standard evaluation sets and online A/B experiments, ensuring the new version is overall better than the old version and preventing the data flywheel from "veering in the wrong direction."
In its mature form, the vast majority of data flywheel operations are automated and encapsulated within a Data / Model Ops platform: from data collection, sample filtering, and annotation task dispatch, to model retraining triggers, evaluation result collection, and rollout decisions — minimizing manual operations and turning model iteration into a stable, controllable engineering process.In its mature form, the vast majority of data flywheel operations are automated and encapsulated within a Data / Model Ops platform: from data collection, sample filtering, and annotation task dispatch, to model retraining triggers, evaluation result collection, and rollout decisions — minimizing manual operations and turning model iteration into a stable, controllable engineering process.
As the number of models and versions grows exponentially, without rigorous lifecycle management, problems such as "models scattered everywhere, chaotic versioning, and difficult rollbacks" easily arise. The goal of ModelOps is to manage models as first-class engineering assets, fully traceable, comparable, and rollback-capable.As the number of models and versions grows exponentially, without rigorous lifecycle management, problems such as "models scattered everywhere, chaotic versioning, and difficult rollbacks" easily arise. The goal of ModelOps is to manage models as first-class engineering assets, fully traceable, comparable, and rollback-capable.
Key points include:Key points include:
Assign a clear version number to each model (e.g., industry-legal-base-v1.2.3) and record:Assign a clear version number to each model (e.g., industry-legal-base-v1.2.3) and record:
Encapsulate the process of "model training complete → automatic evaluation → safety and bias checks → canary release → full rollout" into a CI/CD pipeline.Encapsulate the process of "model training complete → automatic evaluation → safety and bias checks → canary release → full rollout" into a CI/CD pipeline.
In production, multiple model versions often coexist simultaneously (e.g., stable / canary / experimental), compared online through traffic allocation strategies (fixed ratio, user dimension, feature dimension).In production, multiple model versions often coexist simultaneously (e.g., stable / canary / experimental), compared online through traffic allocation strategies (fixed ratio, user dimension, feature dimension).
For industries such as finance, healthcare, and government, every model version change must maintain traceable records: who upgraded which model from which version to which version based on what data and when, and what the post-upgrade impact assessment was. This part typically integrates with the security and compliance infrastructure in section 11.5.For industries such as finance, healthcare, and government, every model version change must maintain traceable records: who upgraded which model from which version to which version based on what data and when, and what the post-upgrade impact assessment was. This part typically integrates with the security and compliance infrastructure in section 11.5.
In engineering implementation, tools such as MLflow / SageMaker / Vertex AI / W&B already provide relatively mature ModelOps capabilities; most enterprises build on top of these with secondary encapsulation tailored to their own processes, constructing a unified internal model registry and release platform.In engineering implementation, tools such as MLflow / SageMaker / Vertex AI / W&B already provide relatively mature ModelOps capabilities; most enterprises build on top of these with secondary encapsulation tailored to their own processes, constructing a unified internal model registry and release platform.
When large models become core business infrastructure, ensuring they are observable, alertable, scalable, and cost-controllable becomes the core responsibility of SRE and platform teams. The monitoring, cost, and reliability layer combines traditional observability systems with large-model-specific metrics, building a multi-dimensional view for operations, algorithms, and management.When large models become core business infrastructure, ensuring they are observable, alertable, scalable, and cost-controllable becomes the core responsibility of SRE and platform teams. The monitoring, cost, and reliability layer combines traditional observability systems with large-model-specific metrics, building a multi-dimensional view for operations, algorithms, and management.
This layer connects monitoring collection, logging / tracing systems on one end, and business KPIs and cost analysis platforms on the other — it is the key pillar ensuring model services are "stable, fast, and cost-effective."This layer connects monitoring collection, logging / tracing systems on one end, and business KPIs and cost analysis platforms on the other — it is the key pillar ensuring model services are "stable, fast, and cost-effective."
In large model systems, traditional CPU / memory / QPS metrics are no longer sufficient; an additional layer of "model-perspective" monitoring is needed to truly see system health. A complete observability system typically includes:In large model systems, traditional CPU / memory / QPS metrics are no longer sufficient; an additional layer of "model-perspective" monitoring is needed to truly see system health. A complete observability system typically includes:
Collect and visualize via Prometheus / Grafana, VictoriaMetrics, etc.:Collect and visualize via Prometheus / Grafana, VictoriaMetrics, etc.:
For large model services, beyond conventional performance metrics, specialized monitoring is also needed:For large model services, beyond conventional performance metrics, specialized monitoring is also needed:
On top of traditional threshold-based alerting, introduce simple statistical monitoring or machine learning models to perform anomaly detection on QPS, latency, error rate, token distribution, etc., automatically alerting when sudden changes occur and triggering self-healing strategies (such as auto-scaling, traffic switching, service degradation).On top of traditional threshold-based alerting, introduce simple statistical monitoring or machine learning models to perform anomaly detection on QPS, latency, error rate, token distribution, etc., automatically alerting when sudden changes occur and triggering self-healing strategies (such as auto-scaling, traffic switching, service degradation).
For algorithm teams, tools such as WhyLabs, Arize, and Evidently AI can also be integrated at this layer to track input distributions, model output characteristics, and drift over the long term, providing signals for subsequent data flywheel and retraining efforts.For algorithm teams, tools such as WhyLabs, Arize, and Evidently AI can also be integrated at this layer to track input distributions, model output characteristics, and drift over the long term, providing signals for subsequent data flywheel and retraining efforts.
One of the most significant operational challenges of large model services is high and volatile costs. Without refined cost analysis and elastic scheduling, it is easy to lose sight of "where the money is being burned" as the business grows, and difficult to make timely adjustments. A mature cost and resource scheduling system typically includes:One of the most significant operational challenges of large model services is high and volatile costs. Without refined cost analysis and elastic scheduling, it is easy to lose sight of "where the money is being burned" as the business grows, and difficult to make timely adjustments. A mature cost and resource scheduling system typically includes:
Use Kubecost, cloud vendor Billing tools, and custom ledgers to break down GPU / CPU / storage / bandwidth costs by dimensions such as model, project, business line, and tenant, so that every team and customer can see their corresponding real resource consumption and expenses.Use Kubecost, cloud vendor Billing tools, and custom ledgers to break down GPU / CPU / storage / bandwidth costs by dimensions such as model, project, business line, and tenant, so that every team and customer can see their corresponding real resource consumption and expenses.
In external API scenarios, this layer also deeply integrates with billing systems, forming a MaaS / API billing and cost accounting platform: billing based on token usage, call count, model specification, and request type, while providing cost and margin analysis for operations / sales.In external API scenarios, this layer also deeply integrates with billing systems, forming a MaaS / API billing and cost accounting platform: billing based on token usage, call count, model specification, and request type, while providing cost and margin analysis for operations / sales.
Once large model capabilities enter highly sensitive industries such as finance, healthcare, and government, security and compliance are no longer "added value" but a prerequisite for entering the scenario. The security, access control, and compliance infrastructure layer is responsible for building system-level defenses spanning access control, data security, privacy protection, and compliance auditing, ensuring model services operate reliably within legal and regulatory frameworks.Once large model capabilities enter highly sensitive industries such as finance, healthcare, and government, security and compliance are no longer "added value" but a prerequisite for entering the scenario. The security, access control, and compliance infrastructure layer is responsible for building system-level defenses spanning access control, data security, privacy protection, and compliance auditing, ensuring model services operate reliably within legal and regulatory frameworks.
This layer connects identity authentication, permission management, key and encryption systems on one end, and model services and logging / auditing platforms on the other — it is the key to turning "a model that can be used" into "a model that dares to be used."This layer connects identity authentication, permission management, key and encryption systems on one end, and model services and logging / auditing platforms on the other — it is the key to turning "a model that can be used" into "a model that dares to be used."
In a large model platform shared by multiple business lines, customers, and roles, without fine-grained access control and tenant isolation, serious problems such as permission abuse, data leakage, and resource contention can easily arise. A comprehensive access and isolation system requires coordination across the following dimensions:In a large model platform shared by multiple business lines, customers, and roles, without fine-grained access control and tenant isolation, serious problems such as permission abuse, data leakage, and resource contention can easily arise. A comprehensive access and isolation system requires coordination across the following dimensions:
Use API Key / Token, OAuth2 / OIDC, enterprise SSO, and other methods for unified identity authentication of internal employees, external partners, and third-party applications. For enterprise users, integrate with existing identity systems (such as AD / LDAP / enterprise IAM) to avoid duplicate account systems.Use API Key / Token, OAuth2 / OIDC, enterprise SSO, and other methods for unified identity authentication of internal employees, external partners, and third-party applications. For enterprise users, integrate with existing identity systems (such as AD / LDAP / enterprise IAM) to avoid duplicate account systems.
Through this layer of mechanisms, the platform can open up large model capabilities to internal and external users while ensuring resource and data security, and provide foundational data for subsequent compliance auditing and issue accountability.Through this layer of mechanisms, the platform can open up large model capabilities to internal and external users while ensuring resource and data security, and provide foundational data for subsequent compliance auditing and issue accountability.
Large models often come into contact with vast amounts of sensitive data (user conversations, business documents, transaction records, etc.). If security or compliance issues arise, the consequences can be extremely severe. Therefore, "multi-layer defense" is needed across the full data lifecycle and the full model invocation chain.Large models often come into contact with vast amounts of sensitive data (user conversations, business documents, transaction records, etc.). If security or compliance issues arise, the consequences can be extremely severe. Therefore, "multi-layer defense" is needed across the full data lifecycle and the full model invocation chain.
This set of capabilities works in concert with the Data / Model Ops and monitoring platforms in sections 11.3 and 11.4, together forming a model operating environment that "can continuously iterate while remaining secure and compliant."This set of capabilities works in concert with the Data / Model Ops and monitoring platforms in sections 11.3 and 11.4, together forming a model operating environment that "can continuously iterate while remaining secure and compliant."
With a complete infrastructure spanning training, inference, security, and operations, an additional "capability layer" facing business and developers is needed — abstracting the underlying large models into components and services that are easier to use and closer to business semantics. This layer is often referred to as the AI platform, application enablement layer, or Copilot platform. Its responsibility is: packaging large models + RAG + Agents + workflows into standardized capabilities, so that business teams and ecosystem partners can rapidly build AI applications.With a complete infrastructure spanning training, inference, security, and operations, an additional "capability layer" facing business and developers is needed — abstracting the underlying large models into components and services that are easier to use and closer to business semantics. This layer is often referred to as the AI platform, application enablement layer, or Copilot platform. Its responsibility is: packaging large models + RAG + Agents + workflows into standardized capabilities, so that business teams and ecosystem partners can rapidly build AI applications.
This layer connects model APIs, RAG engines, and Agent Orchestrators on one end, and business systems such as CRM / ERP / OA / ticketing on the other — it is the key bridge "from model capabilities to business scenarios."This layer connects model APIs, RAG engines, and Agent Orchestrators on one end, and business systems such as CRM / ERP / OA / ticketing on the other — it is the key bridge "from model capabilities to business scenarios."
Compared to early FAQ-style Q&A bots, modern large-model-driven applications are more like "intelligent collaborators that can use tools." The goal of dialogue and Agent orchestration is to upgrade large models from "language generators" to intelligent agents capable of calling tools, executing plans, and coordinating multiple roles.Compared to early FAQ-style Q&A bots, modern large-model-driven applications are more like "intelligent collaborators that can use tools." The goal of dialogue and Agent orchestration is to upgrade large models from "language generators" to intelligent agents capable of calling tools, executing plans, and coordinating multiple roles.
This layer typically leverages existing frameworks such as LangChain, Semantic Kernel, and LlamaIndex, combined with custom Orchestration services, to unify dialogue, tools, workflows, permissions, and auditing within a single "Agent platform."This layer typically leverages existing frameworks such as LangChain, Semantic Kernel, and LlamaIndex, combined with custom Orchestration services, to unify dialogue, tools, workflows, permissions, and auditing within a single "Agent platform."
No matter how powerful a large model is, it cannot naturally master every enterprise's proprietary knowledge, let alone know the latest policies, products, and business rules in real time. RAG + knowledge bases + developer platforms are the key path to connecting this enterprise knowledge, industry knowledge, and real-time data to model capabilities in an engineered way.No matter how powerful a large model is, it cannot naturally master every enterprise's proprietary knowledge, let alone know the latest policies, products, and business rules in real time. RAG + knowledge bases + developer platforms are the key path to connecting this enterprise knowledge, industry knowledge, and real-time data to model capabilities in an engineered way.
This layer ultimately encapsulates complex model and infrastructure capabilities into "reusable, composable business components," helping enterprises — under the premise of security, compliance, and cost control — to truly turn large models into productivity tools that drive business innovation, with lower barriers and faster speed.This layer ultimately encapsulates complex model and infrastructure capabilities into "reusable, composable business components," helping enterprises — under the premise of security, compliance, and cost control — to truly turn large models into productivity tools that drive business innovation, with lower barriers and faster speed.