50-case GPT-as-Policy aligned setGPT-as-Policy 对齐的 50 个 case
10 task families × 5 cases. Each cell: successful cases out of 5, mean native score (0–100) in parentheses where reported. PhysicalRSI publishes task-level success rates only, so its counts are estimates (rate × 5, marked ≈). The best result per row is in bold.
10 个任务族 × 5 个 case。每格为 5 个 case 中的成功数,括号内为原生平均得分(0–100,有报告时列出)。PhysicalRSI 只公开任务级成功率,其成功数为估计值(成功率 × 5,以 ≈ 标记)。每行最佳结果加粗。
Total success rate总成功率
| Task任务 | Auriga | RoboProbe | GPT Direct | GPT Hybrid | PhysicalRSI (est.)PhysicalRSI(估计) |
|---|---|---|---|---|---|
| organize_table | 1 (50) | 0 (45) | 0 (30) | 0 (60) | ≈0 (27) |
| classify_objects_by_language | 4 (88) | 2 (44) | 2 (60) | 1 (38) | ≈0 (21) |
| imitate_sorting_sequence | 4 (83) | 1 (45) | 0 (0) | 2 (53) | ≈0 (1) |
| arrange_largest_number | 5 (100) | 5 (100) | 2 (57) | 2 (50) | ≈1 (21) |
| pack_objects_into_box | 1 (44) | 0 (22) | 1 (50†) | 1 (50) | ≈0 (20) |
| classify_objects | 5 (100) | 2 (54) | 5 (100) | 3 (71) | ≈2 (56) |
| build_tower | 0 (14) | 0 (12) | 0 (12) | 3 (64) | ≈1 (37) |
| make_kong | 0 (0) | 0 (0) | 0 (0) | 2 (40) | ≈4 (82) |
| fold_clothes | 3 (68) | 3 (64) | 2 (40) | 5 (100) | ≈2 (46) |
| put_bottles_into_dustbin | 3 (73) | 1 (43) | 1 (36) | 5 (100) | ≈4 (84) |
| Total合计 | 26/50 (62.0) | 14/50 (42.9) | 13/50 (37.8†) | 24/50 (62.6) | ≈14/50 (39.4) |
The † marker and the PhysicalRSI estimates are explained under “Setup & notes”.
† 标记及 PhysicalRSI 估计值的说明见“设置与说明”。
Auriga mean score per taskAuriga 逐任务平均得分
Mean native score (0–100), judged by the native evaluator.
原生平均得分(0–100),由原生评估器判定。
Overall mean 62.0/100 (conservative over all 50), shown as the dashed line.
总体均值 62.0/100(按全部 50 个 case 保守计算),以虚线表示。
Setup: RoboDojo, Isaac Sim, ARX X5 dual arm · one episode per case · native step limits · success judged by the native evaluator.
设置:RoboDojo、Isaac Sim、ARX X5 双臂 · 每个 case 运行一集 · 使用原生步数上限 · 成功与否由原生评估器判定。
- GPT Direct / Hybrid are the GPT-as-Policy authors' reported results on the same 50 cases.
- GPT Direct / Hybrid 为 GPT-as-Policy 作者在相同 50 个 case 上报告的结果。
- † GPT Direct has no native score for one pack_objects_into_box case; its mean score is over the 48 cases that have one, while its success count is over all 50.
- † GPT Direct 在 pack_objects_into_box 上有 1 个 case 缺少原生得分;其平均得分按有得分的 48 个 case 计算,成功数仍按全部 50 个 case 计算。
- PhysicalRSI: task-level success rate and score from its public technical report and results page, which gives no per-case labels for these 50 cases; its counts are estimated as rate × 5 (marked ≈) and its total as rate × 50.
- PhysicalRSI:任务级成功率与得分取自其公开技术报告与结果页,该来源没有这 50 个 case 的逐 case 标签;成功数按成功率 × 5 估计(以 ≈ 标记),合计按成功率 × 50 估计。
- Other published methods are not listed because their evaluation protocols differ from this 50-case set.
- 其他已发表方法的评估协议与这 50 个 case 不同,因此未列入。
- Auriga: Astra high, 170 model returns per episode; mean native score 62.0/100 (conservative over all 50).
- Auriga:Astra high,每集 170 次模型返回;原生平均得分 62.0/100(按全部 50 个 case 保守计算)。
Harness, Cortex, tools框架、Cortex 与工具
An outer harness wraps the decision core (Cortex, a persistent model session) and the tool layer. All physical motion leaves through a single execution gateway, and every layer writes to the evidence store.
外层框架包裹决策核心(Cortex,一个持续运行的模型会话)和工具层。所有物理运动都经由唯一的执行网关发出,各层都向证据存储写入记录。
Observe
robot_observe returns the RGB frames with the same-frame joint, end-effector and gripper state and an observation ID.
Every later action refers to that observation ID.
观察
robot_observe 返回 RGB 图像,以及同一帧的关节、末端执行器和夹爪状态,并给出观察 ID。
后续动作都引用这个观察 ID。
Plan
A short task plan must be submitted with task_plan before the first physical action.
The plan records the overall intent and is revised when the model's understanding changes.
规划
第一次物理动作之前必须用 task_plan 提交简短的任务计划。
计划记录整体意图,模型的理解发生变化时随之修订。
Preview
robot_action first returns the planned trajectory and its cost without advancing the simulation.
The model checks the preview and adjusts the batch before anything moves.
预览
robot_action 先返回规划好的轨迹及其代价,不推进仿真。
模型检查预览结果,在任何运动发生前调整动作批次。
Execute
The cached plan is executed through the single gateway, which enforces permissions, state checks, budgets and stopping.
Simulation is paused while the model reasons and advances only through actions or explicit waits.
执行
缓存的规划经由唯一的执行网关执行,由网关统一负责权限、状态检查、预算和停止。
模型推理期间仿真暂停,只有动作或显式等待才会推进。
Feedback & evidence
After execution the harness returns images, measured robot state and an execution receipt.
Plans, commands, measured trajectories and scores are recorded separately; on stop the run is scored independently, archived and replayable.
反馈与证据
执行后运行框架返回图像、实测机器人状态和执行回执。
计划、指令、实测轨迹和评分分别记录;停止后独立评分、归档,并可回放。
Four commitments四条原则
The model decides what to do, tools provide the capabilities, and the harness owns the boundary between decision and physical execution.
由模型决定做什么,工具提供能力,框架负责决策与物理执行之间的边界。
- 01
Model-led, tool-assisted模型主导,工具辅助
The model decomposes the task, picks tools and revises its strategy from feedback. Tools expose general capabilities and never ship a task-specific solution.
模型负责拆解任务、选择工具,并根据反馈调整策略。工具只提供通用能力,不内置针对具体任务的解法。
- 02
Closed loop on feedback基于反馈的闭环
Candidate trajectories can be previewed before motion; afterwards the harness returns images and measured robot state, so the model can continue, look again or change plan.
执行前可以预览候选轨迹;执行后框架返回图像和实测的机器人状态,模型据此决定继续、重新观察或修改计划。
- 03
One gateway for every action所有动作经由同一网关
Direct actions, model-written programs and optional policy models all execute through the same entry point, which enforces permissions, state checks, budgets and stopping.
直接动作、模型编写的程序以及可选的策略模型都通过同一个入口执行,由该入口统一负责权限、状态检查、预算和停止。
- 04
Traceable evidence可追溯的证据
Plans, commands, measured trajectories and task scores are recorded separately. Motion finished, program exited and task succeeded stay distinct; failures and unknown states are kept.
计划、指令、实测轨迹和任务得分分别记录。“运动完成”“程序退出”与“任务成功”三者分开记录;失败和未知状态同样保留。
Three ways to put a model on a robot让模型控制机器人的三种方式
The paradigms differ in who decides, what reaches the actuators, and how failures are noticed and repaired.
三种范式的区别在于:由谁决策、什么指令到达执行器,以及如何发现和修复失败。
| (a)Action model (VLA)动作模型(VLA) | (b)General model, direct control通用模型直接控制 | (c) · ours本方法Auriga | |
|---|---|---|---|
| Decides决策 | A learned policy maps images and the instruction to actions.由学习得到的策略将图像和指令映射为动作。 | A general model chooses the next action at each step (e.g. GPT Direct).通用模型在每一步选择下一个动作(如 GPT Direct)。 | A persistent agent session with a task plan, tools and operation memory.一个持续运行的智能体会话,带有任务计划、工具和操作记忆。 |
| Executes执行 | Predicted action chunks go straight to the robot.预测的动作块直接下发给机器人。 | One tool call, one robot step.一次工具调用对应机器人的一步动作。 | Batched single- or dual-arm actions through one gateway: preview, then execute.单臂或双臂的批量动作经由同一网关执行:先预览,后执行。 |
| Feedback反馈 | Implicit, through the next camera frame.隐式反馈,通过下一帧相机图像获得。 | A new observation after each call.每次调用后获得一次新的观察。 | Images, measured joint / gripper state and execution receipts.图像、实测的关节 / 夹爪状态以及执行回执。 |
| Recovery恢复 | Bounded by what the training data covers.受限于训练数据的覆盖范围。 | Re-decided at each call from the new frame; no plan preview or execution receipt.每次调用根据新图像重新决策;没有计划预览和执行回执。 | Re-observe, inspect history, revise the plan; budgets and stopping enforced by the harness.重新观察、查看历史、修订计划;预算与停止由框架强制执行。 |
Selected cases, by capability按能力分组的精选 case
Six cases from the 50-case protocol, grouped by what each one tests. Pick a capability, watch Auriga's episode, then open the side-by-side comparison with a baseline on the same seed and layout.
从 50 case 协议中选出的 6 个 case,按考察的能力分组。选择一项能力,观看 Auriga 的运行,再打开与基线在同一随机种子与布局下的并排对比。
General tools, structured results通用工具,结构化结果
Tools return structured results with images or file references attached as needed. The available set depends on configuration and observation settings.
工具返回结构化结果,并按需附带图像或文件引用。可用的工具集合取决于配置与观察设置。
Click a tool to see what it does, what goes in and out, and an example call.点击工具查看用途、输入输出与调用示例。
core tool核心工具optional, enabled by configuration可选,按配置开启
Run one task运行一个任务
This starts a real model session and a robot simulation. Use a free GPU and a new run directory.
以下命令会启动真实的模型会话和机器人仿真。请使用空闲的 GPU 和新的运行目录。
mamba run -n Auriga python -u -m auriga.gateway \
--benchmark --task general_pickup --seed 0 --layout-id 0 \
--gpu 4 --port 8767 --run-dir runs/my-first-task
# open http://127.0.0.1:8767 for the workbench
# replay an existing run without a model or GPU: --replay --run-dir <run dir>
- Demo videos演示视频six selected cases, Auriga and two baselines on the same seed六个精选 case,Auriga 与两个基线在同一随机种子下对比
- Source code源代码coming soon即将开源
- SimulatorRoboDojo with Isaac Sim installed, in a dedicated environment.仿真器在独立环境中安装 RoboDojo 与 Isaac Sim。
- HardwareA CUDA GPU with free memory for the simulation.硬件一块有足够空闲显存运行仿真的 CUDA GPU。
- Model accessThe native Codex runtime, with authentication and model access configured.模型访问原生 Codex 运行时,并已配置认证与模型访问权限。
Cite this work引用
Until the paper is out, please cite the project page.
论文发布前,请引用本项目页。
@misc{auriga2026,
title = {Auriga: A Model-Led Agent Harness for Embodied Manipulation},
author = {StellarEdge Team},
year = {2026},
howpublished = {Project page},
note = {StellarEdge Team}
}