把 LLM 成本降低 10 倍:先用决策模型做筛选
降低 LLM 成本的做法:用决策模型筛选每一条数据,只把前 5% 交给 LLM 智能体。附可运行的 Swarms 代码,以及实测的成本与耗时。
降低 LLM 成本的做法:用决策模型筛选每一条数据,只把前 5% 交给 LLM 智能体。附可运行的 Swarms 代码,以及实测的成本与耗时。

在高吞吐量的流水线里,降低 LLM 成本最快的办法,是不再把每一条数据都发给 LLM。大多数对销售线索、工单、文档或申请做分诊的流水线,把大部分 token 花在了筛选这一步:逐条阅读,只为判断它值不值得进一步处理。这一步其实是一组带类型的问题,而不是写作任务。本指南会把这一步交给决策模型,也就是 Swarms 的 DecisionModel:它对每一条数据用一次廉价的请求回答一组带类型的问题,而 LLM 智能体只留给少数通过筛选的数据。在我们的实测中,用决策模型筛选的单条成本,比我们能写出的最便宜的 LLM 筛选器低 10.1 倍,响应速度约快 3 倍。
本文的每一个数字,都来自 2026 年 10 月 3 日在 Swarms 16 上运行下面的代码。凡是根据实测样本外推出来的数字,文中都会注明。
这个模式分四个阶段,只有最后一个阶段用到 LLM:
成本之所以能降下来,原因在价格表里。TypeSafe 的 jev-latest 是 DecisionModel 的默认模型,在 Swarms 内置价格表中的价格是每百万输入 token 0.042 美元,输出为 0。gpt-5.4-mini 在 litellm 的价格表中是每百万输入 token 0.75 美元、每百万输出 token 4.50 美元。决策模型不是更小的 LLM,而是另一种返回数字的模型,所以你只为读取数据付费,不为写出内容付费。如果想先了解什么是决策模型、它和 LLM 分类器有什么区别,请阅读什么是决策模型?。本文讲的是如何用它削减账单。
下面的例子为一款数据流水线监控产品筛选主动咨询的销售线索。同样的结构也适用于支持工单、内容审核队列、文档录入和研究资料分诊。仓库里还有一个更长的版本,筛选 500 份合成简历:examples/decision_models/decision_screening_funnel.py。
pip install -U swarms
# or
uv pip install -U swarms示例需要两个密钥:LLM 智能体使用的 OpenAI 密钥,以及决策模型使用的 TypeSafe 密钥。
export OPENAI_API_KEY="sk-..."
export TYPESAFE_API_KEY="..."也可以把两者写进工作目录下的 .env 文件,Swarms 在导入时会自动加载。然后检查版本:
python -c "import swarms; print(swarms.__version__)"DecisionModel 需要 Swarms 16 或更高版本。它的价格表和按请求计价(calculate_cost,由 #2431 加入)包含在已发布到 PyPI 的 16.0.1 中。
在搭建漏斗之前,先看看决策模型返回什么。创建 first_call.py:
from swarms import DecisionModel
# Reads TYPESAFE_API_KEY and calls TypeSafe's jev-latest by default.
model = DecisionModel()
result = model.run(
state={
"ideal_customer": "Data teams at companies with 200 to 5,000 employees",
"lead": {
"employees": 900,
"contact_role": "Head of Data Platform",
"message": "Our dbt models broke three times last month. Can we see a demo next week?",
},
},
questions={
"request": {
"type": "choice",
"instructions": "What is the `lead` asking for?",
"criteria": {
"evaluation": "A demo, a trial, pricing or a vendor evaluation",
"question": "A product or technical question",
"other": "Anything else",
},
},
"fit": {
"type": "score",
"instructions": "How well does the `lead` match the `ideal_customer`?",
"criteria": ["Not a fit", "Partial fit", "Strong fit"],
},
"decision_maker": {
"type": "noul",
"instructions": "The contact leads or manages a data team.",
},
},
)
answers = result["answers"]
print(answers["request"]["choice"], answers["request"]["confidence"])
print(answers["fit"]["score"], answers["fit"]["probabilities"])
print(answers["decision_maker"]["noul"])
print(result["usage"], model.calculate_cost(result["usage"])["total_cost"])python first_call.py我们得到的输出:
evaluation 1.0
1.99 {'0': 0.0, '1': 0.01, '2': 0.99}
0.94
{'input_tokens': 476, 'output_tokens': 70} 1.9992e-05
三个问题,一次请求,千分之二美分。Choice 答案附带置信度,Noul 答案是一个概率,Score 是一个浮点数:1.99 是按概率加权后的等级,正因为如此,分数可以直接排序而不会出现并列。calculate_cost(result["usage"]) 计算一次响应的价格;不带参数调用时,它计算该模型发出的所有请求的总价,漏斗用的就是这种方式。token 与成本追踪指南介绍了智能体和 swarm 上同样的计量方式。
筛选本身就是数据:一份理想客户画像、四个问题和一组权重。把它放在单独的模块里,漏斗脚本和测量脚本就会用完全相同的问题来筛选。创建 leads.py:
import random
IDEAL_CUSTOMER = {
"product": "Monitoring and alerting for data pipelines (Airflow, dbt, Spark)",
"best_fit": [
"Company with 200 to 5,000 employees",
"Runs its own data pipelines in production",
"Contact leads or manages a data or platform team",
"Based in North America or Europe",
],
}
QUESTIONS = {
"request": {
"type": "choice",
"instructions": "What is the `lead` asking for?",
"criteria": {
"evaluation": "A demo, a trial, pricing or a vendor evaluation",
"question": "A product or technical question",
"pitch": "Selling us a service, or looking for a job",
"other": "Students, unsubscribes, spam and anything else",
},
},
"fit": {
"type": "score",
"instructions": "How well does the company in the `lead` match the `ideal_customer`?",
"criteria": [
"Not a fit",
"Weak fit",
"Partial fit",
"Strong fit",
"Ideal customer",
],
},
"intent": {
"type": "noul",
"instructions": "The `lead` describes a concrete pipeline problem or a buying timeline.",
},
"decision_maker": {
"type": "noul",
"instructions": "The contact in the `lead` leads or manages a data or platform team.",
},
}
WEIGHTS = {"fit": 0.45, "intent": 0.30, "decision_maker": 0.25}
MESSAGES = [
("We run about 400 Airflow DAGs and silent failures keep reaching our dashboards. We want a tool in place this quarter.", 3),
("Our dbt models broke three times last month before anyone noticed. Can we see a demo next week?", 3),
("We are evaluating pipeline monitoring vendors for a Q3 rollout across 12 Spark jobs. What does pricing look like at our size?", 3),
("How do you compare to the alerting we already get from Datadog?", 8),
("Do you support Dagster, or only Airflow?", 8),
("Saw your talk at a meetup. Interested, but no timeline yet.", 8),
("I'm a student writing a thesis on data quality. Is there a free license?", 10),
("We are an SEO agency and can get your site to page one of Google.", 12),
("Are you hiring data engineers? My resume is attached.", 10),
("Please remove me from your mailing list.", 10),
("Just browsing.", 12),
]
ROLES = [
"VP of Data",
"Head of Data Platform",
"Data Engineering Manager",
"Senior Data Engineer",
"Analytics Engineer",
"Founder",
"Marketing Coordinator",
"Recruiter",
"Student",
]
INDUSTRIES = [
"Fintech",
"E-commerce",
"Healthcare",
"Logistics",
"Media",
"SaaS",
"Marketing agency",
"University",
]
COUNTRIES = [
"United States",
"Germany",
"United Kingdom",
"Canada",
"Netherlands",
"Brazil",
"India",
"Australia",
]
EMPLOYEES = [8, 25, 60, 150, 400, 900, 2500, 8000, 30000]
def make_leads(count: int, seed: int = 7) -> list:
"""
Generate synthetic inbound leads, most of them poor fits.
Args:
count: Number of leads.
seed: Random seed, so every run screens the same pool.
Returns:
Lead dictionaries.
"""
rng = random.Random(seed)
messages, weights = zip(*MESSAGES)
return [
{
"id": f"L{index:04d}",
"company_industry": rng.choice(INDUSTRIES),
"employees": rng.choice(EMPLOYEES),
"country": rng.choice(COUNTRIES),
"contact_role": rng.choice(ROLES),
"message": rng.choices(messages, weights)[0],
}
for index in range(count)
]
def lead_score(answers: dict) -> float:
"""
Combine the decision model's answers into one score.
Args:
answers: Answers keyed by question id.
Returns:
A score from 0 to 1, or 0 when the lead is not a buyer.
"""
if answers["request"]["choice"] in ("pitch", "other"):
return 0.0
fit_levels = len(QUESTIONS["fit"]["criteria"]) - 1
values = {
"fit": answers["fit"]["score"] / fit_levels,
"intent": answers["intent"]["noul"],
"decision_maker": answers["decision_maker"]["noul"],
}
return sum(WEIGHTS[key] * values[key] for key in WEIGHTS)这个文件里有三个值得照搬的决定。
问题要具体、可核对。 “联系人领导或管理一个数据或平台团队”是一个模型可以对照数据去核实的陈述。“这是一条好线索吗?”则不是;它把标准藏了起来,你之后就没法调整这些标准。
门槛写在代码里,而不是作为一个问题的权重。 一条公司画像完美的供应商推销,得分应该是 0,而不是 0.7。当 Choice 答案是 pitch 或 other 时,lead_score 直接返回 0,只有在其他情况下才应用权重。
权重由你决定。 匹配度占 45%,意向占 30%,职级占 25%。如果销售团队告诉你意向比公司规模更重要,你只需要改一个数字然后重跑。不用改提示词,也不用重新验证输出解析。
make_leads 用固定的随机种子生成一个合成线索池,所以每次运行都筛选同样的 200 条线索;大约七分之一带有明确的购买信息,而同时具备合适公司和合适联系人的要少得多。
现在是漏斗本身。决策模型以每次 20 个并发请求筛选全部 200 条线索,然后由 gpt-5.4 智能体为前 5% 写一份客户简报和一封首次回复。创建 screen.py:
import asyncio
import json
import math
import time
import litellm
from swarms import Agent, DecisionModel, run_agents_with_different_tasks
from leads import IDEAL_CUSTOMER, QUESTIONS, lead_score, make_leads
NUM_LEADS = 200
SHORTLIST_SHARE = 0.05
DEEP_WORK_MODEL = "gpt-5.4"
async def screen_all(model: DecisionModel, leads: list) -> list:
"""
Screen every lead with the decision model, 20 requests at a time.
Args:
model: Decision model that answers the screening questions.
leads: Leads to screen.
Returns:
One result per lead with its score and latency.
"""
limit = asyncio.Semaphore(20)
async def screen(lead: dict) -> dict:
async with limit:
started = time.perf_counter()
response = await model.arun(
state={"ideal_customer": IDEAL_CUSTOMER, "lead": lead},
questions=QUESTIONS,
)
return {
"lead": lead,
"score": lead_score(response["answers"]),
"seconds": time.perf_counter() - started,
}
return await asyncio.gather(*(screen(lead) for lead in leads))
def llm_cost(model_name: str, usage: dict) -> float:
"""
Price an agent's token usage with litellm's price table.
Args:
model_name: Model the agent ran on.
usage: The agent's usage dictionary.
Returns:
Cost in US dollars.
"""
prompt_cost, completion_cost = litellm.cost_per_token(
model=model_name,
prompt_tokens=usage["input_tokens"],
completion_tokens=usage["output_tokens"],
)
return prompt_cost + completion_cost
leads = make_leads(NUM_LEADS)
model = DecisionModel()
started = time.perf_counter()
screened = asyncio.run(screen_all(model, leads))
screen_wall = time.perf_counter() - started
screen_cost = model.calculate_cost()
shortlist_size = math.ceil(NUM_LEADS * SHORTLIST_SHARE)
shortlist = sorted(screened, key=lambda r: r["score"], reverse=True)[
:shortlist_size
]
print(f"Screened {NUM_LEADS} leads in {screen_wall:.1f} s")
print(
f" {screen_cost['input_tokens']} input tokens, "
f"${screen_cost['total_cost']:.5f} total, "
f"${screen_cost['total_cost'] / NUM_LEADS:.7f} per lead"
)
print(
f" {sum(r['seconds'] for r in screened) / NUM_LEADS:.2f} s per request"
)
print(f"\nTop 5 of the {shortlist_size}-lead shortlist:")
for result in shortlist[:5]:
lead = result["lead"]
print(
f" {result['score']:.2f} {lead['contact_role']}, "
f"{lead['employees']} employees, {lead['country']}: "
f"{lead['message'][:50]}..."
)
writers = [
Agent(
agent_name=f"Account-Executive-{r['lead']['id']}",
system_prompt=(
"You are an account executive. For the lead you are given, write "
"a three-bullet account brief (why they fit, what to ask, what "
"could block the deal), then a short first reply that refers to "
"their message."
),
model_name=DEEP_WORK_MODEL,
max_loops=1,
output_type="final",
print_on=False,
)
for r in shortlist
]
started = time.perf_counter()
briefs = run_agents_with_different_tasks(
[
(
writer,
f"Ideal customer: {json.dumps(IDEAL_CUSTOMER)}\n\n"
f"Lead: {json.dumps(r['lead'])}",
)
for writer, r in zip(writers, shortlist)
]
)
deep_wall = time.perf_counter() - started
deep_cost = sum(llm_cost(DEEP_WORK_MODEL, w.usage) for w in writers)
print(
f"\nDeep work on {shortlist_size} leads with {DEEP_WORK_MODEL}: "
f"{deep_wall:.1f} s, ${deep_cost:.4f}, "
f"${deep_cost / shortlist_size:.4f} per lead"
)
print(
f"Pipeline total: ${screen_cost['total_cost'] + deep_cost:.4f} "
f"for {NUM_LEADS} leads"
)
print(f"\nBrief for {shortlist[0]['lead']['id']}:\n\n{briefs[0]}")python screen.py我们的运行结果(简报在第一条要点之后截断):
Screened 200 leads in 5.3 s
127791 input tokens, $0.00537 total, $0.0000268 per lead
0.46 s per request
Top 5 of the 10-lead shortlist:
0.98 VP of Data, 400 employees, United States: We run about 400 Airflow DAGs and silent failures ...
0.98 VP of Data, 2500 employees, United Kingdom: We are evaluating pipeline monitoring vendors for ...
0.95 Data Engineering Manager, 2500 employees, United States: We run about 400 Airflow DAGs and silent failures ...
0.88 Senior Data Engineer, 400 employees, Germany: Our dbt models broke three times last month before...
0.88 Senior Data Engineer, 900 employees, Germany: We are evaluating pipeline monitoring vendors for ...
Deep work on 10 leads with gpt-5.4: 6.4 s, $0.0473, $0.0047 per lead
Pipeline total: $0.0527 for 200 leads
Brief for L0082:
- **Why they fit:** Strong ICP match: 400 employees, US-based fintech, and the contact is a **VP of Data**. They run a meaningful production footprint with **~400 Airflow DAGs**, and the pain is exactly what we solve: **silent pipeline failures reaching dashboards**.
这份入围名单正是销售负责人会亲手挑出来的:具体的流水线痛点、明确的采购时间表、数据团队的负责人、规模合适的公司。筛选全部 200 条线索只花了半美分。有几个实现细节很重要:
arun。 DecisionModel.arun 是 run 的异步版本。信号量把并发上限设为 20,客户端会对限流和过载(408、429、500、502、503、504、529)进行带退避的重试,并遵守 retry-after。output_type="final" 只返回答案,print_on=False 让日志里不出现面板,而独立的智能体可以避免一条线索的上下文混进另一条。run_agents_with_different_tasks 会并发运行它们。model.calculate_cost() 把 provider 报告的 token 数在全部 200 次请求上求和;每个智能体的 usage 则用 litellm 的价格表计价。成本倍数是否可信,取决于它的基线。公平的比较不是“决策模型对比一个写长文的智能体”,而是决策模型对比你真正会上线的最便宜的 LLM 筛选器:gpt-5.4-mini、简短的系统提示词、同样的四个问题、只输出 JSON。我们还测量了团队经常会走到的那个变体:让筛选器为每个答案附上一句理由。创建 compare.py:
import json
import random
import time
import litellm
from swarms import Agent, DecisionModel
from leads import IDEAL_CUSTOMER, QUESTIONS, make_leads
SAMPLE_SIZE = 3
SCREEN_MODEL = "gpt-5.4-mini"
JSON_PROMPT = (
"You screen inbound sales leads against an ideal customer profile. "
"Reply with JSON only: "
'{"request": "evaluation" | "question" | "pitch" | "other", '
'"fit": 0-4 (0 not a fit, 4 ideal customer), '
'"intent": 0-1 (the lead describes a concrete pipeline problem or buying timeline), '
'"decision_maker": 0-1 (the contact leads or manages a data or platform team)}'
)
SCREENERS = {
"LLM, JSON only": JSON_PROMPT,
"LLM, JSON with reasons": JSON_PROMPT
+ ' Add a "reasons" field with one sentence per answer.',
}
sample = random.Random(1).sample(make_leads(200), SAMPLE_SIZE)
results = {}
model = DecisionModel()
seconds = []
for lead in sample:
started = time.perf_counter()
model.run(
state={"ideal_customer": IDEAL_CUSTOMER, "lead": lead},
questions=QUESTIONS,
)
seconds.append(time.perf_counter() - started)
results[f"Decision model ({model.model_name})"] = (
model.calculate_cost()["total_cost"],
sum(seconds),
)
for name, system_prompt in SCREENERS.items():
cost, seconds = 0.0, 0.0
for lead in sample:
# A fresh agent per lead keeps earlier leads out of the context.
screener = Agent(
agent_name="Screener",
system_prompt=system_prompt,
model_name=SCREEN_MODEL,
max_loops=1,
output_type="final",
print_on=False,
)
started = time.perf_counter()
screener.run(
f"Ideal customer: {json.dumps(IDEAL_CUSTOMER)}\n\n"
f"Lead: {json.dumps(lead)}"
)
seconds += time.perf_counter() - started
prompt_cost, completion_cost = litellm.cost_per_token(
model=SCREEN_MODEL,
prompt_tokens=screener.usage["input_tokens"],
completion_tokens=screener.usage["output_tokens"],
)
cost += prompt_cost + completion_cost
results[f"{name} ({SCREEN_MODEL})"] = (cost, seconds)
baseline_cost = results[f"Decision model ({model.model_name})"][0]
print(f"Per lead, measured on {SAMPLE_SIZE} leads:")
for name, (cost, seconds) in results.items():
per_lead = cost / SAMPLE_SIZE
print(
f" {name:<40} ${per_lead:.7f} {seconds / SAMPLE_SIZE:.2f} s "
f"{cost / baseline_cost:5.1f}x cost "
f"${per_lead * 100_000:,.2f} per 100k leads"
)python compare.py我们的运行结果:
Per lead, measured on 3 leads:
Decision model (jev-latest) $0.0000269 0.20 s 1.0x cost $2.69 per 100k leads
LLM, JSON only (gpt-5.4-mini) $0.0002710 0.66 s 10.1x cost $27.10 per 100k leads
LLM, JSON with reasons (gpt-5.4-mini) $0.0006168 0.91 s 22.9x cost $61.67 per 100k leads
三条线索是一个很小的样本,这是为了让测试保持便宜;想要更精确的估计,就调大 SAMPLE_SIZE。有两点说明这些数字是稳定的。决策模型在样本上的单条成本(0.0000269 美元)与 200 条线索那次运行(0.0000268 美元)一致;而 LLM 筛选器的输入是固定的提示词加一条简短的线索,JSON 输出的长度也几乎不变。这里的延迟是顺序测得的,一次一个请求;第 3 步中每次请求 0.46 秒,是决策模型在 20 路并发下的数字。
倍数来自输出。只输出 JSON 的筛选器,答案只有几十个 token,但 gpt-5.4-mini 的输出 token 价格是输入 token 的六倍,而它的输入价格本身就已经是决策模型的约 18 倍。让筛选器解释自己的判断,差距就会扩大到两倍以上,达到 22.9 倍。
下面是整条流水线在我们实际运行的规模下的成本,以及外推到 100,000 条线索的结果。实测数字来自上面的运行。外推数字是把实测的单条成本乘以数据条数。
| 流水线 | 200 条线索 | 100,000 条线索(外推) |
|---|---|---|
决策模型筛选全部,gpt-5.4 智能体处理前 5% | $0.0527(实测) | $2.69 + $23.65 = $26.34 |
gpt-5.4-mini 筛选全部(JSON),智能体处理前 5% | $0.0542(外推)+ $0.0473(实测)= $0.1015 | $27.10 + $23.65 = $50.75 |
不筛选:gpt-5.4 智能体处理每一条线索 | $0.95(由 10 条外推) | $473 |
这张表可以从三个方向读。
仅看筛选:便宜 10.1 倍。 这是决策模型所替代的阶段,也是本文标题里的倍数。如果对比会写理由的筛选器,则是 22.9 倍。
整条漏斗:比不做漏斗便宜约 18 倍。 把每条线索都交给 gpt-5.4 智能体,200 条线索要花 0.95 美元;漏斗只花了 0.0527 美元。大部分节省来自截断到 5%,而决策模型让这次截断几乎不花钱:筛选只占漏斗账单的 10%。
筛选变便宜之后,深度工作就成了主要开销。 在我们的运行中,10 次智能体调用花了 0.0473 美元,200 次筛选请求花了 0.0054 美元。LLM 成本优化的下一个杠杆不再是筛选,而是 SHORTLIST_SHARE:入围比例每降低一个百分点,省下的钱就超过整个筛选阶段的花费。
耗时也呈现同样的形态。决策模型顺序处理时每条线索 0.20 秒,gpt-5.4-mini 筛选器则是 0.66 秒;筛选全部 200 条线索的实际耗时为 5.3 秒。
本文的实际测试,包括上面所有运行,总共花费约 0.06 美元。
这个模式的前提是:数据量大,值得深度处理的只占一小部分,而且问题能写得清楚。只要其中任何一条不成立,用它的理由也就不成立了。
if 语句里,既免费又准确。决策模型应该用在需要阅读才能做出的判断上:意向、语气、相关性、从职位名称推断职级。我们在本文中也没有用标注数据测量筛选的准确率;上面的入围名单只是一次合理性检查,而不是评估。在替换生产环境中的筛选之前,请在几百条你已经标注过的数据上同时运行两种筛选器,比较各自的入围名单。
漏斗只是更大一类结构中的一种形态;智能体编排模式介绍了其他形态,而决策模型可以作为路由器或评判者嵌入其中的大多数。随 DecisionModel 一起发布的完整内容,包括可直接替换的 Cloudflare Clef provider,以及 examples/decision_models/ 中的其他示例,请参阅 Swarms v16「Overclock」发布说明。
找到那个让 LLM 逐条阅读数据、只为做出一个判断的步骤,把它换成一个以带类型的值回答同样问题的决策模型。通过筛选的数据仍然交给 LLM。然后直接检查质量:在切换之前,用两种筛选器处理已标注的数据,比较它们的入围名单。
在 Swarms 内置价格表中,jev-latest 的价格是每百万输入 token 0.042 美元,输出为 0。我们的线索加上四个问题,平均每次请求约 640 个输入 token,花费 0.0000268 美元。DecisionModel.calculate_cost() 会根据 provider 报告的 usage 返回精确数字。
LLM 分类器生成一段文字,你再把它解析成标签。决策模型直接返回标签,附带概率和置信度,并且不对输出收费。什么是决策模型?对两者的区别有深入的解释。
可以。DecisionModel(model_name="clef") 或 model_name="clef-flash" 会路由到 Cloudflare Workers AI,需要 CLOUDFLARE_ACCOUNT_ID 和 CLOUDFLARE_AUTH_TOKEN。内置价格分别是每百万输入 token 0.24 美元和 0.09 美元。我们在本文中没有运行 Clef,所以上面的测量结果只适用于 jev-latest。
这取决于你的问题和你的数据,没有任何公开基准能替你回答。具体、可核对的问题比什么都管用。在你自己的已标注数据上测量,设置一个置信度阈值,把没把握的答案交给人工或智能体,并抽查被筛选拒绝的数据。
有问题或反馈?欢迎加入我们的 Discord 社区,或查阅文档。

什么是决策模型?TypeSafe Jev 与 Cloudflare Clef 如何用经过校准的置信度回答带类型的问题,而不是生成文本,以及如何在 Swarms 中使用它们。

用 Swarms MCPDeployer 在 Python 中把 AI 智能体变成 MCP 服务器:API key、自定义认证、token 校验器、把 swarm 作为工具,以及一个调用它的客户端智能体。

用 Swarms 在 Python 中实现 tree of thoughts:求解 24 点游戏,对比 BFS 与 DFS、propose 与 sample、value 与 vote,控制成本,并查看真实的调用次数。