OCTER.AI / RESEARCH

Octer-Decision: A System One decision model that proposes better options

Technical research on high-frequency open-choice inference, decision-making, and recursive self-improvement.

Read the research
In this article 5 sections

Abstract摘要

At Octer, we developed Octer-Decision for high-frequency decisions whose best action may lie outside the supplied candidate set. Our Choice head scores the listed actions and a learned coverage sentinel; we invoke a generation head only when that sentinel wins. We use outcome-linked traces to recursively improve both model weights and the decision harness, validating each version before release. In our advertising evaluation, Octer-Decision + RSI has approximately half the P95 decision-path latency of Jev (0.22 versus 0.44 s) and improves the profit index by 18.4%.

我们在 Octer 研发了 Octer-Decision,用于最优动作可能位于预设候选集之外的高频决策。Choice 头同时评估既有动作和一个学习得到的覆盖缺口特殊项;只有该项获胜时,我们才激活生成头。我们将决策轨迹与结果关联,用于递归更新模型权重和决策 harness,并在发布前验证每个新版本。在我们的广告评估中,Octer-Decision + RSI 的 P95 决策链路延迟约为 Jev 的一半(0.22 对 0.44 秒),利润指数提高 18.4%。

01.The fixed-candidate bottleneck固定候选集的决策瓶颈

Fast decision models can return a probability distribution over a supplied set of actions. Jev made this structured, probabilistic interface prominent; our team had already put Octer-Decision to work before Jev's public announcement. The limitation we address is the candidate set itself: when it omits a suitable response, its highest-scoring member is still drawn from an incomplete set.

高速决策模型能够针对给定动作集合输出概率分布。Jev 让结构化概率决策受到关注;在 Jev 公开发布前,我们已经将 Octer-Decision 投入使用。我们要解决的是候选集合本身的覆盖问题:当集合缺少合适动作时,最高分仍来自一个不完整的集合。

We therefore make candidate-set coverage part of inference. A learned sentinel lets our Choice head flag an inadequate action set. Routine requests take a single-step probability path; coverage gaps activate conditional generation of one structured candidate. We then use delayed outcomes to revise both the model and the harness that defines, composes, and validates actions.

因此,我们把候选集合的覆盖程度纳入推断。学习得到的特殊项让 Choice 头识别现有动作覆盖不足:常规请求走单步概率判断路径,覆盖缺口才触发一个结构化候选的条件生成。随后,我们利用延迟反馈同时修订模型,以及定义、组合、校验动作的 harness。

02.Conditional generation and recursive updates条件生成与递归更新

We send Octer-Decision a point-in-time state, a typed action set, and execution constraints. Around the model, we built the components that construct that state, schedule requests, compose compatible actions, enforce constraints, and retain versioned outcome traces.

我们向 Octer-Decision 输入决策时状态、类型化动作集合和执行约束。围绕模型,我们构建了状态编译、请求调度、动作组合、约束执行和带版本的结果轨迹等组件。

What we added to the decision layer

我们在决策层增加的组件

Component组件System role系统职责
State / Outcome CompilerReconstructs the state available at decision time and marks delayed observations.还原决策时可知的状态,并标记延迟到达的观测结果。
Decision FabricSchedules concurrent decisions and preserves versioned state–action traces.调度并发判断,保留带版本的状态—动作轨迹。
Policy ComposerCombines local choices into a globally feasible, typed action set.将局部判断组合为全局可执行的类型化动作集合。
Safety & OutcomeApplies permissions and hard constraints; links delayed outcomes to each version.执行权限与硬约束,并将延迟结果关联到对应版本。

A choice model with a conditional open branch

带条件开放分支的 Choice 模型

We use a shared encoder and Choice head to produce calibrated probabilities over the supplied actions and NCA (“no correct action”). When NCA has the highest probability, our conditional autoregressive decoder proposes a new action; otherwise inference ends after the Choice head. We train in two stages: RLCD calibration of the encoder–Choice path, then decoder-only supervised fine-tuning while that path is frozen. New candidates undergo utility assessment and constraint checks before execution; in later RSI cycles, we can revise the next version after outcome-based validation.

我们用共享编码器与 Choice 头对给定动作及 NCA(“现有选项均不正确”)输出校准后的概率。若 NCA 概率最高,我们才让条件自回归解码器生成新动作;否则推断在 Choice 头结束。训练分两阶段:先以 RLCD 校准编码器与 Choice 路径,再冻结该路径,仅对解码器做监督微调。新候选经过价值评估与约束校验后进入执行;在后续 RSI 周期中,我们根据结果验证修订下一版本。

Figure 1. Octer-Decision's two-stage training and conditional inference path. NCA = no correct action; the orange decoder path runs only when NCA is selected. Click to enlarge.
图 1。Octer-Decision 的两阶段训练与条件推断路径。NCA 表示“现有选项均不正确”;仅选中 NCA 时才运行橙色解码路径。点击可放大。

The RSI feedback loop

RSI 如何形成闭环

We record the observed state, candidate set, selected action, model and harness versions, safety intervention, and delayed outcome in each trace. Our RSI loop proposes updates to both model weights and the harness; we promote a version only after offline evaluation and a guarded release.

我们在每条轨迹中记录观测状态、候选集合、最终动作、模型与 harness 版本、安全拦截及延迟到达的结果。我们的 RSI 闭环同时提出模型权重与 harness 的更新;新版本经过离线评估和受控发布后才进入下一轮。

03.A high-frequency decision case study高频决策实验案例

3.1 Experimental protocol

3.1 实验协议

Our evaluation protocol assigns 300 campaigns to three arms (100 per arm): Jev, Octer-Decision, and Octer-Decision + RSI. We hold campaign eligibility, per-round spend allocation, attribution rules, platform write frequency, and safety limits constant across arms. Eight decision rounds, followed by a conversion-maturity hold, cover 6.048 million decision opportunities across 15 targets per campaign. Octer-Decision keeps its initial model and harness; Octer-Decision + RSI receives validated updates to both after rounds 2, 4, and 6. We compare the fixed version with the continuously improving mechanism over the full evaluation.

我们的评估协议将 300 个广告活动随机分为三组(每组 100 个):Jev、Octer-Decision 和 Octer-Decision + RSI。我们让三组使用相同的活动资格、逐轮花费分配、归因规则、平台写入频率和安全约束。八轮决策及随后等待转化成熟的阶段,覆盖每个活动 15 个投放目标上的 604.8 万次决策机会。Octer-Decision 沿用初始模型与 harness;Octer-Decision + RSI 在第 2、4、6 轮后接收经验证的模型和 harness 更新。我们比较固定版本与持续自进化机制在完整评测过程中的表现。

Primary outcome:主指标: profit index = 100 × group net profit per unit of spend / Jev net profit per unit of spend.利润指数 = 100 × 各组单位广告花费净利润 / Jev 组单位广告花费净利润。

We use campaigns as the unit of randomization and inference, with campaign-clustered bootstrap resampling for 95% intervals. We measure latency and throughput in the same deployment envelope under a 128-concurrent-request load sample.

我们以活动为随机化与统计推断单位,按活动聚类 Bootstrap 计算 95% 区间。延迟与吞吐在同一实验部署中、128 并发请求的负载样本下测量。

Recording 1. An early interface recording of campaign selection, data analysis, and target-level bid decisions. It illustrates one class of advertising action; the experiment uses the protocol above. Click to enlarge.
录屏 1。早期界面中的活动选择、数据分析与目标级竞价判断。竞价是广告动作的一类;实验结果按上文协议统计。点击可原地放大。

3.2 Profit across eight rounds

3.2 八轮收益轨迹

We set the Jev arm's full-experiment profit per unit of ad spend to a baseline of 100. Across eight rounds, Octer-Decision + RSI reaches 118.4, compared with 110.4 for Octer-Decision and 100.0 for Jev. The difference is 18.4 index points over Jev (campaign-clustered 95% interval 10.8–25.9) and 8.0 points over Octer-Decision (0.9–15.1).

我们将 Jev 组的全实验单位广告花费净利润设为基准 100。八轮实验中,Octer-Decision + RSI 达到 118.4,Octer-Decision 为 110.4,Jev 为 100.0。相对 Jev 的差值为 18.4 个指数点(按活动聚类的 95% 区间:10.8–25.9);相对 Octer-Decision 的差值为 8.0 个指数点(0.9–15.1)。

Figure 2 · Profit index by round
图 2 · 各轮利润指数
Same randomized campaigns throughout; orange dashed lines mark Octer-Decision + RSI releases after rounds 2, 4, and 6.
同一组随机分配的活动贯穿全程;橙色虚线标记 Octer-Decision + RSI 在第 2、4、6 轮后的更新。
JevOcter-DecisionOcter-Decision + RSI Profit index (Jev = 100)利润指数(Jev = 100)8090100110120130140 12345678Experiment round实验轮次 134.9112.3101.9
Figure 2. Each point is that round's matured outcome under equal spend; the eight-round arm estimates appear in the text. Release markers locate model-and-harness updates, while the randomized Octer-Decision arm supports the update comparison.
图 2。各点为等额花费下该轮成熟归因后的结果;八轮组均值见正文。发布标记定位模型与 harness 更新,更新效果通过随机分配的 Octer-Decision 组比较。

3.3 Profit, latency, and decision capacity

3.3 收益、延迟与决策容量

We plot the three arms by P95 decision-path latency and the eight-round profit index; circle area encodes relative peak throughput under the shared 128-concurrent-request load. Octer-Decision + RSI lies above and to the left of Jev: a profit index of 118.4 versus 100.0, 0.22 s versus 0.44 s, and 69% higher throughput.

我们把三组同时放在 P95 决策链路延迟与八轮利润指数坐标上,圆点面积表示共用 128 并发负载下的相对峰值吞吐。Octer-Decision + RSI 位于 Jev 的左上方:利润指数 118.4 对 100.0,延迟 0.22 对 0.44 秒,吞吐提高 69%。

Figure 3 · Joint profit–latency–throughput result
图 3 · 收益—延迟—吞吐联合结果
X = P95 decision-path latency in seconds (lower is better); Y = profit index (Jev = 100; higher is better); circle area = relative throughput; vertical line = 95% profit interval.
横轴为决策链路 P95 延迟(秒,越低越好);纵轴为利润指数(Jev = 100,越高越好);圆点面积为相对吞吐;竖线为利润 95% 区间。
Profit index (Jev = 100)利润指数(Jev = 100)8090100110120130 00.10.20.30.40.5P95 decision-path latency (s)决策链路延迟 P95(秒) JevOcter-DecisionOcter-Decision + RSI
Figure 3. Each point summarizes the same randomized arm (100 campaigns). Profit intervals use campaign-clustered bootstrap estimates; P95 and throughput include state preparation, model inference, policy composition, and safety checks. The occasional generated-action route is included for Octer-Decision + RSI.
图 3。每个点概括同一随机实验组(100 个活动)。利润区间按活动聚类 Bootstrap 计算;P95 与吞吐包括状态准备、模型推断、策略组合及安全检查。Octer-Decision + RSI 的延迟统计包含按需生成新动作的路径。

3.4 Open-choice behavior in the same experiment

3.4 同一实验中的开放候选行为

We also track when Octer-Decision + RSI uses its open path. The Choice head selects an existing action in 95.8% of decision windows and activates candidate generation in 4.2%. Our rule checks find 94.1% of generated candidates feasible; execution still passes through the shared safety gate. The controller writes to the ad platform in 8.6% of windows, with no executed hard-limit violation recorded in this 300-campaign experiment.

我们还记录了 Octer-Decision + RSI 何时进入开放路径。Choice 头在 95.8% 的窗口直接选中既有动作,4.2% 的窗口触发新候选生成。规则校验判定 94.1% 的生成候选具备可执行性;实际执行仍须通过共用安全门。8.6% 的窗口产生广告平台写入,这项 300 活动实验未记录执行级硬约束违规。

04.Recovery under physical perturbations物理扰动下的恢复决策

We build our mobile pick-and-place recovery evaluation on ManiSkill, with six perturbation classes and a recovery protocol developed by the Octer team. The reported metrics come from our closed-loop evaluation runs. An LLM supplies the task plan; Octer-Decision selects the next skill after a disturbance. When an occupied destination makes the listed placement actions infeasible, its open branch can propose a temporary placement followed by repositioning. The controller checks reachability and collision constraints before executing the proposed skill sequence.

我们基于 ManiSkill 构建移动取放任务的中断恢复评测,由 Octer 团队设计六类扰动与恢复协议,报告指标来自我们记录的闭环评测轨迹。LLM 提供任务计划,Octer-Decision 在扰动后决定下一项技能。例如,目标位置被占用导致既有放置动作失效时,开放分支可以提出“先临时放置,再调整位置”的技能组合,由控制器校验可达性与碰撞约束后执行。

We evaluate all three models on 240 paired initial states per disturbance, holding perception, skill library, safety checks, and compute budgets constant. RSI uses separate development trajectories; evaluation objects and layouts remain held out. Recovery success requires autonomous completion without unsafe contact; takeover rate counts episodes requiring human assistance. P95 measures state-ready to validated-action latency, including conditional generation; an independent controller handles immediate stopping. The lower three rows stress situations where the supplied candidates are insufficient.

我们在每类扰动的 240 个配对初始状态上评测三组模型,各组共用感知、技能库、安全校验和计算预算。RSI 使用独立开发轨迹,评估物体与布局保持留出。恢复成功要求自主完成任务且无不安全接触;接管率统计需要人工介入的回合。P95 从状态就绪计至动作通过校验,包含条件生成;即时停车由独立控制器负责。图中后三类场景重点检验给定候选不足时的恢复能力。

Figure 4 · Recovery across six perturbation classes
图 4 · 六类扰动下的恢复表现
Six robot disturbances and three models: recovery success, P95 decision latency, and human takeover rate. 六类机器人扰动、三组模型的恢复成功率、P95 决策延迟和人工接管率。
Figure 4. Octer's perturbation-recovery evaluation built on ManiSkill. Each row compares the same disturbance across models; each model has three metric columns. A shared color scale applies within each metric. Success and takeover rates are calculated from our evaluation episode counts.
图 4。Octer 基于 ManiSkill 构建的扰动恢复评测。每行对应同一类扰动,每个模型包含三列指标;相同指标共用色阶。成功率与接管率由本组评测回合计数计算。

Under combined perturbations, Octer-Decision + RSI reaches 81.7% recovery success versus Jev's 52.5%, a gain of 29.2 percentage points; takeover falls from 32.1% to 12.5%. Its P95 increases with task complexity, from 71 ms on aisle obstruction to 182 ms on combined perturbations. Within this evaluation, the model retains a latency advantage over Jev's 286 ms while completing more recoveries autonomously.

在复合扰动下,Octer-Decision + RSI 的恢复成功率为 81.7%,相比 Jev 的 52.5% 提高 29.2 个百分点;接管率从 32.1% 降至 12.5%。其 P95 随任务复杂度上升,从通道受阻时的 71 毫秒增至复合扰动时的 182 毫秒。在本组评测中,模型完成更多自主恢复,同时保持相对 Jev 286 毫秒的延迟优势。

05.Conclusion总结

Octer-Decision connects fast probabilistic choice, conditional proposal of new options, and recursive improvement of model weights and harness. The Choice path keeps routine decisions short; the open branch expands the available actions when coverage is insufficient. Outcome-linked trajectories let RSI improve both how the model chooses and what it can choose next. Together, these mechanisms target timely, adaptable decisions as the environment changes.

Octer-Decision 将高速概率选择、按需提出新选项,以及模型权重与 harness 的递归更新结合起来。Choice 路径缩短常规决策,开放分支在候选覆盖不足时扩展可选动作,RSI 则通过结果轨迹持续改进选择能力与动作空间。这三项机制共同服务于一个目标:让模型在环境变化中持续做出及时、有效的决策。

Octer-Decision Choice head and conditional decoder architecture enlarged Ad-bidding recording enlarged
OCTER.AI / DECISION

Try Octer-Decision

Tell us a little about your use case. Our team will contact you about a free trial.

Applicant type

We will only use these details to follow up on your request.