Abstract摘要
At Octer, we developed Octer-Decision for high-frequency decisions whose best action may lie outside the supplied candidate set. Our Choice head scores the listed actions and a learned coverage sentinel; we invoke a generation head only when that sentinel wins. We use outcome-linked traces to recursively improve both model weights and the decision harness, validating each version before release. In our advertising evaluation, Octer-Decision + RSI has approximately half the P95 decision-path latency of Jev (0.22 versus 0.44 s) and improves the profit index by 18.4%.
我们在 Octer 研发了 Octer-Decision,用于最优动作可能位于预设候选集之外的高频决策。Choice 头同时评估既有动作和一个学习得到的覆盖缺口特殊项;只有该项获胜时,我们才激活生成头。我们将决策轨迹与结果关联,用于递归更新模型权重和决策 harness,并在发布前验证每个新版本。在我们的广告评估中,Octer-Decision + RSI 的 P95 决策链路延迟约为 Jev 的一半(0.22 对 0.44 秒),利润指数提高 18.4%。
01.The fixed-candidate bottleneck固定候选集的决策瓶颈
Fast decision models can return a probability distribution over a supplied set of actions. Jev made this structured, probabilistic interface prominent; our team had already put Octer-Decision to work before Jev's public announcement. The limitation we address is the candidate set itself: when it omits a suitable response, its highest-scoring member is still drawn from an incomplete set.
高速决策模型能够针对给定动作集合输出概率分布。Jev 让结构化概率决策受到关注;在 Jev 公开发布前,我们已经将 Octer-Decision 投入使用。我们要解决的是候选集合本身的覆盖问题:当集合缺少合适动作时,最高分仍来自一个不完整的集合。
We therefore make candidate-set coverage part of inference. A learned sentinel lets our Choice head flag an inadequate action set. Routine requests take a single-step probability path; coverage gaps activate conditional generation of one structured candidate. We then use delayed outcomes to revise both the model and the harness that defines, composes, and validates actions.
因此,我们把候选集合的覆盖程度纳入推断。学习得到的特殊项让 Choice 头识别现有动作覆盖不足:常规请求走单步概率判断路径,覆盖缺口才触发一个结构化候选的条件生成。随后,我们利用延迟反馈同时修订模型,以及定义、组合、校验动作的 harness。
02.Conditional generation and recursive updates条件生成与递归更新
We send Octer-Decision a point-in-time state, a typed action set, and execution constraints. Around the model, we built the components that construct that state, schedule requests, compose compatible actions, enforce constraints, and retain versioned outcome traces.
我们向 Octer-Decision 输入决策时状态、类型化动作集合和执行约束。围绕模型,我们构建了状态编译、请求调度、动作组合、约束执行和带版本的结果轨迹等组件。
What we added to the decision layer
我们在决策层增加的组件
| Component | 组件 | System role | 系统职责 |
|---|---|---|---|
| State / Outcome Compiler | Reconstructs the state available at decision time and marks delayed observations. | 还原决策时可知的状态,并标记延迟到达的观测结果。 | |
| Decision Fabric | Schedules concurrent decisions and preserves versioned state–action traces. | 调度并发判断,保留带版本的状态—动作轨迹。 | |
| Policy Composer | Combines local choices into a globally feasible, typed action set. | 将局部判断组合为全局可执行的类型化动作集合。 | |
| Safety & Outcome | Applies permissions and hard constraints; links delayed outcomes to each version. | 执行权限与硬约束,并将延迟结果关联到对应版本。 |
A choice model with a conditional open branch
带条件开放分支的 Choice 模型
We use a shared encoder and Choice head to produce calibrated probabilities over the supplied actions and NCA (“no correct action”). When NCA has the highest probability, our conditional autoregressive decoder proposes a new action; otherwise inference ends after the Choice head. We train in two stages: RLCD calibration of the encoder–Choice path, then decoder-only supervised fine-tuning while that path is frozen. New candidates undergo utility assessment and constraint checks before execution; in later RSI cycles, we can revise the next version after outcome-based validation.
我们用共享编码器与 Choice 头对给定动作及 NCA(“现有选项均不正确”)输出校准后的概率。若 NCA 概率最高,我们才让条件自回归解码器生成新动作;否则推断在 Choice 头结束。训练分两阶段:先以 RLCD 校准编码器与 Choice 路径,再冻结该路径,仅对解码器做监督微调。新候选经过价值评估与约束校验后进入执行;在后续 RSI 周期中,我们根据结果验证修订下一版本。
The RSI feedback loop
RSI 如何形成闭环
We record the observed state, candidate set, selected action, model and harness versions, safety intervention, and delayed outcome in each trace. Our RSI loop proposes updates to both model weights and the harness; we promote a version only after offline evaluation and a guarded release.
我们在每条轨迹中记录观测状态、候选集合、最终动作、模型与 harness 版本、安全拦截及延迟到达的结果。我们的 RSI 闭环同时提出模型权重与 harness 的更新;新版本经过离线评估和受控发布后才进入下一轮。
03.A high-frequency decision case study高频决策实验案例
3.1 Experimental protocol
3.1 实验协议
Our evaluation protocol assigns 300 campaigns to three arms (100 per arm): Jev, Octer-Decision, and Octer-Decision + RSI. We hold campaign eligibility, per-round spend allocation, attribution rules, platform write frequency, and safety limits constant across arms. Eight decision rounds, followed by a conversion-maturity hold, cover 6.048 million decision opportunities across 15 targets per campaign. Octer-Decision keeps its initial model and harness; Octer-Decision + RSI receives validated updates to both after rounds 2, 4, and 6. We compare the fixed version with the continuously improving mechanism over the full evaluation.
我们的评估协议将 300 个广告活动随机分为三组(每组 100 个):Jev、Octer-Decision 和 Octer-Decision + RSI。我们让三组使用相同的活动资格、逐轮花费分配、归因规则、平台写入频率和安全约束。八轮决策及随后等待转化成熟的阶段,覆盖每个活动 15 个投放目标上的 604.8 万次决策机会。Octer-Decision 沿用初始模型与 harness;Octer-Decision + RSI 在第 2、4、6 轮后接收经验证的模型和 harness 更新。我们比较固定版本与持续自进化机制在完整评测过程中的表现。
We use campaigns as the unit of randomization and inference, with campaign-clustered bootstrap resampling for 95% intervals. We measure latency and throughput in the same deployment envelope under a 128-concurrent-request load sample.
我们以活动为随机化与统计推断单位,按活动聚类 Bootstrap 计算 95% 区间。延迟与吞吐在同一实验部署中、128 并发请求的负载样本下测量。
3.2 Profit across eight rounds
3.2 八轮收益轨迹
We set the Jev arm's full-experiment profit per unit of ad spend to a baseline of 100. Across eight rounds, Octer-Decision + RSI reaches 118.4, compared with 110.4 for Octer-Decision and 100.0 for Jev. The difference is 18.4 index points over Jev (campaign-clustered 95% interval 10.8–25.9) and 8.0 points over Octer-Decision (0.9–15.1).
我们将 Jev 组的全实验单位广告花费净利润设为基准 100。八轮实验中,Octer-Decision + RSI 达到 118.4,Octer-Decision 为 110.4,Jev 为 100.0。相对 Jev 的差值为 18.4 个指数点(按活动聚类的 95% 区间:10.8–25.9);相对 Octer-Decision 的差值为 8.0 个指数点(0.9–15.1)。
3.3 Profit, latency, and decision capacity
3.3 收益、延迟与决策容量
We plot the three arms by P95 decision-path latency and the eight-round profit index; circle area encodes relative peak throughput under the shared 128-concurrent-request load. Octer-Decision + RSI lies above and to the left of Jev: a profit index of 118.4 versus 100.0, 0.22 s versus 0.44 s, and 69% higher throughput.
我们把三组同时放在 P95 决策链路延迟与八轮利润指数坐标上,圆点面积表示共用 128 并发负载下的相对峰值吞吐。Octer-Decision + RSI 位于 Jev 的左上方:利润指数 118.4 对 100.0,延迟 0.22 对 0.44 秒,吞吐提高 69%。
3.4 Open-choice behavior in the same experiment
3.4 同一实验中的开放候选行为
We also track when Octer-Decision + RSI uses its open path. The Choice head selects an existing action in 95.8% of decision windows and activates candidate generation in 4.2%. Our rule checks find 94.1% of generated candidates feasible; execution still passes through the shared safety gate. The controller writes to the ad platform in 8.6% of windows, with no executed hard-limit violation recorded in this 300-campaign experiment.
我们还记录了 Octer-Decision + RSI 何时进入开放路径。Choice 头在 95.8% 的窗口直接选中既有动作,4.2% 的窗口触发新候选生成。规则校验判定 94.1% 的生成候选具备可执行性;实际执行仍须通过共用安全门。8.6% 的窗口产生广告平台写入,这项 300 活动实验未记录执行级硬约束违规。
04.Recovery under physical perturbations物理扰动下的恢复决策
We build our mobile pick-and-place recovery evaluation on ManiSkill, with six perturbation classes and a recovery protocol developed by the Octer team. The reported metrics come from our closed-loop evaluation runs. An LLM supplies the task plan; Octer-Decision selects the next skill after a disturbance. When an occupied destination makes the listed placement actions infeasible, its open branch can propose a temporary placement followed by repositioning. The controller checks reachability and collision constraints before executing the proposed skill sequence.
我们基于 ManiSkill 构建移动取放任务的中断恢复评测,由 Octer 团队设计六类扰动与恢复协议,报告指标来自我们记录的闭环评测轨迹。LLM 提供任务计划,Octer-Decision 在扰动后决定下一项技能。例如,目标位置被占用导致既有放置动作失效时,开放分支可以提出“先临时放置,再调整位置”的技能组合,由控制器校验可达性与碰撞约束后执行。
We evaluate all three models on 240 paired initial states per disturbance, holding perception, skill library, safety checks, and compute budgets constant. RSI uses separate development trajectories; evaluation objects and layouts remain held out. Recovery success requires autonomous completion without unsafe contact; takeover rate counts episodes requiring human assistance. P95 measures state-ready to validated-action latency, including conditional generation; an independent controller handles immediate stopping. The lower three rows stress situations where the supplied candidates are insufficient.
我们在每类扰动的 240 个配对初始状态上评测三组模型,各组共用感知、技能库、安全校验和计算预算。RSI 使用独立开发轨迹,评估物体与布局保持留出。恢复成功要求自主完成任务且无不安全接触;接管率统计需要人工介入的回合。P95 从状态就绪计至动作通过校验,包含条件生成;即时停车由独立控制器负责。图中后三类场景重点检验给定候选不足时的恢复能力。
Under combined perturbations, Octer-Decision + RSI reaches 81.7% recovery success versus Jev's 52.5%, a gain of 29.2 percentage points; takeover falls from 32.1% to 12.5%. Its P95 increases with task complexity, from 71 ms on aisle obstruction to 182 ms on combined perturbations. Within this evaluation, the model retains a latency advantage over Jev's 286 ms while completing more recoveries autonomously.
在复合扰动下,Octer-Decision + RSI 的恢复成功率为 81.7%,相比 Jev 的 52.5% 提高 29.2 个百分点;接管率从 32.1% 降至 12.5%。其 P95 随任务复杂度上升,从通道受阻时的 71 毫秒增至复合扰动时的 182 毫秒。在本组评测中,模型完成更多自主恢复,同时保持相对 Jev 286 毫秒的延迟优势。
05.Conclusion总结
Octer-Decision connects fast probabilistic choice, conditional proposal of new options, and recursive improvement of model weights and harness. The Choice path keeps routine decisions short; the open branch expands the available actions when coverage is insufficient. Outcome-linked trajectories let RSI improve both how the model chooses and what it can choose next. Together, these mechanisms target timely, adaptable decisions as the environment changes.
Octer-Decision 将高速概率选择、按需提出新选项,以及模型权重与 harness 的递归更新结合起来。Choice 路径缩短常规决策,开放分支在候选覆盖不足时扩展可选动作,RSI 则通过结果轨迹持续改进选择能力与动作空间。这三项机制共同服务于一个目标:让模型在环境变化中持续做出及时、有效的决策。