一、什么是结构化输出?

LangChain 的结构化输出(Structured Output)指的是:

要求模型最终返回一个符合预定义结构的数据对象(比如固定字段的 JSON、Pydantic 模型、TypedDict),而不再是无格式的自然语言文本。

核心目标:把"自然语言回答"变成"程序可以稳定消费的数据"。

举个例子:

不是让模型输出:

盗梦空间在2010年上映,导演是克里斯托弗·诺兰,评分9.3。

而是让它输出成这样的结构:

{
    "title": "盗梦空间",
    "year": 2010,
    "director": "克里斯托弗·诺兰",
    "rating": 9.3
}

价值有三个方面:

  1. 更容易被代码处理:下游系统可以直接读字段,不用再从自然语言里做解析。

  2. 结果更稳定:减少"模型说法变了但意思差不多"导致的解析失败。

  3. 更适合工程化:适用于表单抽取、分类、路由、工具参数生成、工作流状态传递等场景。


二、传统方式 vs 结构化输出

传统方式(繁琐、不推荐)

# 1. 提示词要求 JSON
prompt = "以JSON格式返回:{name, age, occupation}"
response = model.invoke(prompt)
​
# 2. 手动解析
import json
data = json.loads(response.content)
​
# 3. 手动验证类型
if not isinstance(data['age'], int):
    raise ValueError("age must be int")
​
# 4. 手动创建对象
person = Person(**data)

结构化输出(一步到位)

structured_llm = model.with_structured_output(Person)
person = structured_llm.invoke("张三是一名 30 岁的软件工程师")
# ✅ 自动解析、验证、创建对象

为什么结构化输出这么受欢迎?

在没有 Pydantic 等结构化方案之前,开发者需要写大量 Prompt 苦口婆心地求模型"请返回 JSON,不要带任何解释",然后自己写繁琐的 json.loads() 和 try...except。

有了 Pydantic 等方案结合 .with_structured_output() 之后:

  • Prompt 变干净了:字段的 description 直接充当了 Prompt 的一部分。

  • 类型安全:编辑器能自动补全,运行前就能做类型检查。

  • 极其稳定:依托模型厂商底层的 JSON 模式,输出错误率降到极低。


三、四种模式总览

LangChain 1.x 支持多种 Schema 与结构化输出方式:

模式

特点

返回类型

Pydantic

字段校验、描述、嵌套结构,功能最丰富

Schema 类实例

TypedDict

轻量类型约束

字典

JSON Schema

与前后端/跨语言接口最通用

字典

dataclass

简洁的数据类定义

字典

关键区别:

  • 只有 Pydantic 返回的是 Schema 类实例,其余三种都返回字典。

  • 只有 Pydantic 在类型不匹配时会抛出异常(强校验)。

模型支持情况:大部分现代模型支持(OpenAI gpt-4 / gpt-3.5-turbo、Anthropic claude-3、Groq llama-3 等,通过函数调用)。某些旧模型不支持,此时 LangChain 会回退到"提示词 + JSON 解析"。


四、模式一:Pydantic(生产首选)

Pydantic 在运行时强制执行类型提示,确保数据正确性和一致性,是生产场景首选。

4.1 基本使用三要素

  1. 所有结构化输出的数据模型都必须继承 BaseModel。

  2. 使用类型提示:str、int、float、List[xxx]、Optional[xxx] 等。

  3. 用 Field() 添加字段默认值和描述,帮助 LLM 理解字段含义。

完整示例:

# 1. 模型初始化
from langchain.chat_models import init_chat_model
from dotenv import load_dotenv
import os
​
load_dotenv(override=True)
​
CLOSEAI_API_KEY = os.getenv("CLOSEAI_API_KEY")
CLOSEAI_BASE_URL = os.getenv("CLOSEAI_BASE_URL")
​
model = init_chat_model(
    model="gpt-5.4-mini",
    model_provider="openai",
    api_key=CLOSEAI_API_KEY,
    base_url=CLOSEAI_BASE_URL
)
​
# 2. 定义 Pydantic 模型
from pydantic import BaseModel, Field
​
class Person(BaseModel):
    """人物信息"""
    name: str = Field(description="姓名")
    age: int = Field(description="年龄")
    occupation: str = Field(description="职业")
​
# 3. 使用 with_structured_output
structured_llm = model.with_structured_output(Person)
​
result = structured_llm.invoke("张三是一名 30 岁的软件工程师")
​
print(result)
print(type(result))
# name='张三' age=30 occupation='软件工程师'
# <class '__main__.Person'>
​
print(result.name)        # "张三"
print(result.age)         # 30
print(result.occupation)  # "软件工程师"

⚠️ 没有描述,LLM 可能格式错误。 所以一定要写 description。

情感分析示例:

class SentimentAnalysis(BaseModel):
    """情感分析结果"""
    sentiment: str = Field(description="情感倾向:positive/negative/neutral")
    confidence: float = Field(description="置信度,0-1之间")
    keywords: list[str] = Field(description="关键词列表")
​
structured_model = model.with_structured_output(SentimentAnalysis)
​
text = "这个课程内容很实用,学到了很多知识,强烈推荐!"
result = structured_model.invoke(f"分析以下文本的情感:\n{text}")
​
print(result.sentiment)   # positive
print(result.confidence)  # 0.99
print(result.keywords)    # ['实用', '学到了很多知识', '强烈推荐']

4.2 高级特性

① 可选字段(Optional)

LLM 未填充某些字段怎么办?用 Optional 指定字段可选。

from typing import Optional
​
class Person(BaseModel):
    """人物信息"""
    name: str = Field(description="姓名")
    age: Optional[int] = Field(description="年龄")   # 可选
    occupation: str = Field(description="职业")
  • 不用 Optional:Person(name='张三', age=0, occupation='医生')(会被填 0)

  • 用 Optional:Person(name='张三', age=None, occupation='医生')(填 None)

② 默认值

格式:Field(default="默认值", description="描述")

class Product(BaseModel):
    """产品信息"""
    name: str = Field(description="产品名称")
    price: float = Field(description="价格")
    description: Optional[str] = Field(description="产品描述")
    stock: int = Field(default=100, description="库存")   # 默认 100

⚠️ 注意:不同模型提供商对 default 字段的支持是不同的,测试时要以实际平台为准。

③ 枚举类型(Enum / Literal)

用枚举限制字段可选值。

from enum import Enum
from typing import Optional
from pydantic import BaseModel, Field
​
class Priority(str, Enum):
    LOW = "低"
    MEDIUM = "中"
    HIGH = "高"
​
class CustomerInfo(BaseModel):
    """客户信息"""
    name: str = Field(description="客户姓名")
    phone: str = Field(description="电话号码")
    email: Optional[str] = Field(description="邮箱")
    issue: str = Field(description="问题描述")
    urgency: Priority = Field(description="紧急程度")   # 只能是 低/中/高

或者用 Literal 直接写死(更简单):

from typing import Literal
​
class CustomerInfo(BaseModel):
    """客户信息"""
    name: str = Field(description="客户姓名")
    urgency: Literal["低", "中", "高"] = Field(description="紧急程度")

应用场景:自动填充 CRM、工单自动分类、客服辅助。

④ 列表提取(List)

from typing import List
​
class Person(BaseModel):
    name: str
    age: int
​
class PersonList(BaseModel):
    """人物列表信息"""
    people: List[Person]    # 多个 Person 对象
​
structured_llm = model.with_structured_output(PersonList)
result = structured_llm.invoke("张三 30岁,李四 25岁")
# people=[Person(name='张三', age=30), Person(name='李四', age=25)]

应用场景:批量处理用户评论、自动生成分析报告、发现产品改进点、自动化财务处理、OCR 后结构化。

⑤ 嵌套结构

from pydantic import BaseModel, Field
from typing import List
​
class Actor(BaseModel):
    """演员信息"""
    name: str = Field(description="演员姓名")
    role: str = Field(description="饰演的角色")
​
class Movie(BaseModel):
    """电影信息"""
    title: str = Field(description="电影标题")
    year: int = Field(description="上映年份")
    director: str = Field(description="导演")
    cast: List[Actor] = Field(description="演员列表")     # 嵌套列表
    rating: float = Field(description="评分")
​
structured_model = model.with_structured_output(Movie)
response = structured_model.invoke("请介绍电影《盗梦空间》")
​
print(response.cast)  # [Actor(name='莱昂纳多·迪卡普里奥', role='柯布'), ...]

⚠️ LLM 能力有限,复杂嵌套结构可能会出错。 建议:

  • 嵌套层级 ≤ 3 层(4 层以上容易出错)

  • 使用清晰的 description

  • 必要时拆分成多个调用

⑥ 限制条件(参数校验)

用 Field 的约束参数做字段级校验:

from pydantic import ValidationError
​
class User(BaseModel):
    name: str = Field(min_length=2, max_length=20)
    age: int = Field(ge=0, le=150)   # 0~150
    email: str
​
try:
    user = User(name="李四", age=200, email="li@example.com")
except ValidationError as e:
    print(e.errors()[0]['msg'])
    # Input should be less than or equal to 150

常用约束:min_length、max_length、ge(≥)、le(≤)、gt(>)。


五、模式二:TypedDict(轻量)

TypedDict 是 Python 3.8+ 引入的类型提示工具,即"带类型声明的字典结构"。适合快速定义字典结构、无需 Pydantic 重量级功能的场景。

普通 dict 没有类型信息,TypedDict 可以声明字段和类型。但它主要是类型声明,不是运行时强校验器。

from typing_extensions import TypedDict
​
class MovieDict(TypedDict):
    title: str
    year: int
    director: str
    rating: float
​
movie: MovieDict = {
    "title1": "盗梦空间",   # 字段名不一致,IDE 会标记,但运行不报错
    "year": 2010,
    "director": "克里斯托弗·诺兰",
    "rating": 8.8,
}

注意:字段名写错(如 title1)IDE 静态检查会提示,但不会导致运行时异常。

用 Annotated 附加描述

Annotated 在"类型"之外附加元数据(类似 Pydantic 的 Field)。

from typing_extensions import TypedDict, Annotated
from typing import List
​
class Actor(TypedDict):
    """演员情况"""
    name: Annotated[str, "演员姓名"]
    role: Annotated[str, "饰演的角色"]
​
class Movie(TypedDict):
    """电影情况"""
    title: Annotated[str, "电影标题"]
    year: Annotated[int, "上映年份"]
    director: Annotated[str, "导演"]
    cast: Annotated[List[Actor], "演员列表"]      # 嵌套列表
    rating: Annotated[float, "评分"]
​
structured_llm = model_with_closeai.with_structured_output(Movie)
resp = structured_llm.invoke("给我介绍下电影《盗梦空间》")
​
print(resp['title'])   # 盗梦空间(返回的是字典,用 [] 访问)

六、模式三:JSON Schema(跨语言通用)

JSON Schema 与前/后端、跨语言接口最通用,直接传符合 JSON Schema 标准的字典或字符串。

json_schema = {
    "title": "Movie",
    "description": "A movie with details",
    "type": "object",
    "properties": {
        "title": {"type": "string", "description": "The title of the movie"},
        "year": {"type": "integer", "description": "The year the movie was released"},
        "director": {"type": "string", "description": "The director of the movie"},
        "rating": {"type": "number", "description": "The movie's rating out of 10"}
    },
    "required": ["title", "year", "director", "rating"]
}
​
structured_model = model.with_structured_output(
    json_schema,
    method="json_schema"
)

返回的是字典,不校验字段匹配。


七、模式四:dataclass

用标准库的 @dataclass 装饰器定义数据结构,让"数据结构定义"更简洁、清晰。

from dataclasses import dataclass
from pydantic import Field
​
@dataclass
class Movie:
    """
    电影的详细信息
    """
    title: str = Field(description="电影标题")
    year: int = Field(description="电影上映年份")
    director: str = Field(description="导演")
    rating: float = Field(description="电影评分,满分十分")
​
structured_model = model.with_structured_output(Movie)
response = structured_model.invoke("给出盗梦空间的信息")
print(response)       # {'title': '盗梦空间', 'year': 2010, ...}
print(type(response)) # <class 'dict'>

⚠️ 注意:@dataclass 修饰的类是数据类,被标准库标记并携带字段元信息,可以作为 LangChain 的 Schema;而手写 __init__ 等方法的普通类不能替代。另外它返回的是未经校验的字典。


八、四种模式的关键区别:类型校验

通过一个"fake server"(模拟 DeepSeek 服务端,故意返回字段不匹配的数据)可以直观地看到差异。

假设模型返回了字段名不匹配的数据(title1、year2),各自的处理如下:

模式

字段不匹配时的表现

Pydantic

❌ 抛出 ValidationError(Field required)

TypedDict

✅ 按字典输出,不报错

JSON Schema

✅ 按字典输出,不报错

dataclass

✅ 按字典输出,不报错

Pydantic 的报错示例:

ValidationError: 2 validation errors for MovieModel
title
  Field required [type=missing, ...]
year
  Field required [type=missing, ...]

结论:用 Pydantic 定义 schema,收到响应后会强制校验,字段不匹配则抛异常;其余三种方式不校验。

这也是为什么生产环境更推荐 Pydantic —— 它能第一时间帮你发现"模型吐错了字段"的问题。


九、获取结构化结果的方式

除了四种 Schema 模式,获取结构化结果本身也有两种方式。

方式 1:with_structured_output(推荐)

最新、最简洁的 API,直接让模型"理解"数据结构并返回解析好的对象。

还可以传 include_raw=True 参数,返回解析前的原始 AIMessage,从而访问令牌用量等元数据:

from pydantic import BaseModel, Field
from rich import print as rprint
​
class Movie(BaseModel):
    """电影信息"""
    title: str = Field(description="电影标题")
    year: int = Field(description="上映年份")
    director: str = Field(description="导演")
    rating: float = Field(description="评分(10分制)")
​
model_with_structure = model.with_structured_output(Movie, include_raw=True)
resp = model_with_structure.invoke("给我介绍下电影《星际穿越》")
​
print(type(resp))   # <class 'dict'>

返回结果是一个字典,包含三个字段:

  • raw:返回的原始 AIMessage(含 token 用量等元数据)。

  • parsed:解析后的输出(比如 Movie(...) 实例)。

  • parsing_error:解析错误。当前用 Pydantic,格式不符合 schema 会报错;其余三种方式不会。

方式 2:输出解析器(传统,不推荐)

更传统的方法:在提示词里明确指示模型输出特定格式文本,再用解析器转换。

流程:提示词指导 → 模型生成文本 → 解析器转换。

from langchain_core.output_parsers import JsonOutputParser
from langchain_core.prompts import ChatPromptTemplate
from pydantic import BaseModel, Field
​
# 1. 提示词模板
prompt_template = ChatPromptTemplate.from_messages([
    ("system", "回答用户问题,必须始终输出一个包含title(电影标题)和year(上映年份)的 JSON 对象"),
    ("human", "问题:{question}")
])
​
# 2. 定义结构
class Movie(BaseModel):
    """电影信息"""
    title: str = Field(description="电影标题")
    year: int = Field(description="上映年份")
​
# 3. 创建解析器
parser = JsonOutputParser(pydantic_object=Movie)
​
# 4. 创建链(用 | 管道拼接)
chain = prompt_template | model | parser
​
# 5. 调用
response = chain.invoke({"question": "介绍电影《盗梦空间》"})
print(response)   # {'title': '盗梦空间', 'year': 2010}

这种方式的缺点是:依赖 Prompt"求"模型输出 JSON,稳定性不如 with_structured_output。


十、工作原理图解

以 Pydantic + with_structured_output 为例,内部流程分为四步:

第 1 步:定义结构

from pydantic import BaseModel, Field
​
class BookInfo(BaseModel):
    title: str = Field(description="书名")
    author: str = Field(description="作者名字")
    tags: list[str] = Field(description="书籍的标签或分类")

第 2 步:协议转换

LangChain 内部调用 Pydantic 的底层方法(如 model_json_schema()),把你写的 Python 代码自动转成标准 JSON Schema——详细描述有哪些字段、字段类型(string、array 等)以及字段描述。

第 3 步:模型交互与强约束

LangChain 把 JSON Schema 包装进模型 API 请求中。

  • 现代方法(.with_structured_output):现代大模型普遍支持"函数/工具调用"或"JSON Mode",LangChain 把 JSON Schema 作为 Tools 传入。

  • 大模型侧约束:像 OpenAI 的 strict=True 参数会启动语法采样约束(Grammar-based sampling),模型解码生成 token 时严格按 JSON Schema 的语法树选择,从底层保证输出格式不走样。

第 4 步:自动解析与验证

模型返回 JSON 字符串后,PydanticStructuredOutputParser(解析器)接管:

  1. 解析(Parsing):把字符串解析为 Python 字典。

  2. 验证(Validation):把字典喂给 Pydantic 模型,自动检查类型。缺必填字段或类型错误会抛验证错误(或触发 LangChain 重试)。

  3. 返回(Return):通过验证后,你拿到的是可直接点出属性的 Pydantic 对象(如 result.title)。


十一、小结

主题

核心要点

什么是结构化输出

让模型返回预定义结构的数据对象,而非自然语言文本

三种价值

易处理、更稳定、适合工程化

四种 Schema

Pydantic(首选)/ TypedDict / JSON Schema / dataclass

关键区别

仅 Pydantic 返回类实例 + 强校验;其余返回字典、不校验

Pydantic 高级特性

Optional、默认值、Enum/Literal、List、嵌套(≤3层)、约束

获取方式

with_structured_output(推荐)/ 输出解析器(传统)

include_raw=True

返回 raw/parsed/parsing_error 三字段

工作原理

定义结构→转 JSON Schema→模型强约束→解析验证返回