Ollaya:在 NVIDIA RTX 4090 上通过 HTTP API 处理一个包含五个问题的请求的中位延迟(laya 使用 fp16,其余使用 fp32)。Jev:第三方基准测试(AbdelStark/jev-benchmarks、nibzard/decision-model-benchmark)中托管 API 的请求中位延迟,其中包含网络开销。由于测试环境不同,请将其视为数量级上的对比。
Ollaya 提供 /v1/systemone 和 /v1/models 两个端点,请求与响应格式与 TypeSafe 完全一致。官方 TypeSafe Python SDK 0.7.1 无需任何修改即可连接本地服务器。
`# Point the TypeSafe SDK at Ollaya export TYPESAFE_BASE_URL=http://localhost:11435 export TYPESAFE_API_KEY=local # any value works export TYPESAFE_DEFAULT_MODEL=laya
…or call the compatible endpoint directly
curl http://localhost:11435/v1/systemone -d ‘{ “model”: “laya”, “state”: “Can I get an invoice for last month?”, “questions”: { “intent”: { “type”: “choice”, “instructions”: “What does the customer want?”, “criteria”: { “invoice”: “Needs an invoice or receipt”, “refund”: “Wants money back”, “other”: “Anything else” } } } }'`
响应
{ "model": "laya:en", "answers": { "intent": { "type": "choice", "choice": "invoice", "confidence": 0.9547, "probabilities": { "invoice": 0.9698, "refund": 0.0172, "other": 0.013 } } }, "usage": { "input_tokens": 43, "output_tokens": 0 } }
首先可以使用来自 Convai Innovations 的 Laya 系列模型:包括一个英语模型、一个支持 100 多种语言的模型、一个针对结构化决策微调的模型,以及一个可替你自动选择的模型路由器。
更多开放决策模型已在计划中:von,以及通过 llama.cpp 运行的基于 GGUF LLM 的决策模型。
工单、邮件和用户消息往往是你手中最敏感的数据。有了 Ollaya,这些数据可以在它们本来所在的位置就地完成评分。
本地运行
借助 ONNX Runtime 在你自己的机器上运行,支持 CPU 或 NVIDIA GPU。服务器默认监听 127.0.0.1。
开放权重
权重来自各模型作者的 Hugging Face 仓库,锁定到特定 commit 并通过 sha256 校验。Ollaya 从不重新托管这些权重,运行时采用 Apache-2.0 许可证。
无按 token 计费
只要硬件扛得住,想跑多少决策就跑多少。没有计量,也没有 API 账单。
已校准
输出的是可以直接设定阈值的概率。经过温度拟合后,Laya 的校准误差(ECE)为 0.081,而 Jev 为 0.246。