Auriga

A model-led agent harness for embodied manipulation.

面向具身操作的模型主导智能体框架。

A persistent model session understands the task, calls tools and controls the robot, revising its plan from real execution feedback. Every physical action passes through one execution gateway, and every run leaves a replayable evidence trail.

一个持续运行的模型会话负责理解任务、调用工具并控制机器人,并根据真实的执行反馈修订计划。所有物理动作都经过同一个执行网关,每次运行都会留下可回放的证据记录。

StellarEdge Team

Code代码 coming soon即将开源 Demo演示 videos视频
26 / 50tasks solved个 case 成功 62.0mean score平均得分 10task families类任务
background: Auriga on arrange_largest_number, head camera背景:Auriga 执行 arrange_largest_number,头部相机
Results结果

50-case GPT-as-Policy aligned setGPT-as-Policy 对齐的 50 个 case

10 task families × 5 cases. Each cell: successful cases out of 5, mean native score (0–100) in parentheses where reported. PhysicalRSI publishes task-level success rates only, so its counts are estimates (rate × 5, marked ≈). The best result per row is in bold.

10 个任务族 × 5 个 case。每格为 5 个 case 中的成功数,括号内为原生平均得分(0–100,有报告时列出)。PhysicalRSI 只公开任务级成功率,其成功数为估计值(成功率 × 5,以 ≈ 标记)。每行最佳结果加粗。

Total success rate总成功率

Aurigaother methods其他方法
Total success rate over 50 cases: Auriga 52.0%, GPT Hybrid 48.0%, PhysicalRSI 28.9% (task-level rate), RoboProbe 28.0%, GPT Direct 26.0% 0% 20% 40% 60% Auriga 52.0% GPT Hybrid 48.0% PhysicalRSI 28.9% RoboProbe 28.0% GPT Direct 26.0%
Successful cases / 50; PhysicalRSI is its reported task-level success rate (estimated counts in the table).成功 case 数 / 50;PhysicalRSI 为其报告的任务级成功率(表中为估计成功数)。
26/50Auriga successful cases (52%)Auriga 成功的 case(52%)
62.0mean native score / 100, conservative over all 50原生平均得分 / 100,按全部 50 个 case 保守计算
170model returns per episode (Astra high)每集模型返回次数(Astra high)
Task任务AurigaRoboProbeGPT DirectGPT HybridPhysicalRSI (est.)PhysicalRSI(估计)
organize_table1 (50)0 (45)0 (30)0 (60)≈0 (27)
classify_objects_by_language4 (88)2 (44)2 (60)1 (38)≈0 (21)
imitate_sorting_sequence4 (83)1 (45)0 (0)2 (53)≈0 (1)
arrange_largest_number5 (100)5 (100)2 (57)2 (50)≈1 (21)
pack_objects_into_box1 (44)0 (22)1 (50†)1 (50)≈0 (20)
classify_objects5 (100)2 (54)5 (100)3 (71)≈2 (56)
build_tower0 (14)0 (12)0 (12)3 (64)≈1 (37)
make_kong0 (0)0 (0)0 (0)2 (40)≈4 (82)
fold_clothes3 (68)3 (64)2 (40)5 (100)≈2 (46)
put_bottles_into_dustbin3 (73)1 (43)1 (36)5 (100)≈4 (84)
Total合计26/50 (62.0)14/50 (42.9)13/50 (37.8†)24/50 (62.6)≈14/50 (39.4)

The † marker and the PhysicalRSI estimates are explained under “Setup & notes”.

† 标记及 PhysicalRSI 估计值的说明见“设置与说明”。

Architecture架构

Harness, Cortex, tools框架、Cortex 与工具

An outer harness wraps the decision core (Cortex, a persistent model session) and the tool layer. All physical motion leaves through a single execution gateway, and every layer writes to the evidence store.

外层框架包裹决策核心(Cortex,一个持续运行的模型会话)和工具层。所有物理运动都经由唯一的执行网关发出,各层都向证据存储写入记录。

Architecture: harness containing Cortex, tool layer, execution gateway and evidence store; RoboDojo and web workbench outside HARNESS · RUN CONTROL框架 · 运行控制 Task instruction / user任务指令 / 用户 Cortex persistent model session · understand, plan, decide持续运行的模型会话 · 理解、规划、决策 Tool layer工具层 observe · geometry · action · program · memory观察 · 几何 · 动作 · 程序 · 记忆 Unified execution gateway统一执行网关 validate · plan · execute · feedback校验 · 规划 · 执行 · 反馈 tool calls / results工具调用 / 结果 action request / execution feedback动作请求 / 执行反馈 Run records & evidence运行记录与证据 plans · commands · measured trajectories · images · scores计划 · 指令 · 实测轨迹 · 图像 · 得分 RoboDojo robot & environment机器人与环境 Isaac Sim · ARX X5 dual armIsaac Sim · ARX X5 双臂 robot action / observation机器人动作 / 观察 Web workbench / experiment libraryWeb 工作台 / 实验库 live view · terminal · replay实时画面 · 终端 · 回放 user control用户控制
Solid arrows carry requests and results; dashed lines are evidence writes. Model-written programs and optional policy models request motion through the same gateway.实线箭头表示请求与结果;虚线表示证据写入。模型编写的程序与可选的策略模型也通过同一网关请求运动。
Run loop: observe, plan, preview, execute, feedback, then continue or stop ObservePlanPreview ExecuteFeedback Done? continue / handle failure adjust stop → scoring, archive, replay
STEP 01

Observe

robot_observe returns the RGB frames with the same-frame joint, end-effector and gripper state and an observation ID.

Every later action refers to that observation ID.

步骤 01

观察

robot_observe 返回 RGB 图像,以及同一帧的关节、末端执行器和夹爪状态,并给出观察 ID。

后续动作都引用这个观察 ID。

STEP 02

Plan

A short task plan must be submitted with task_plan before the first physical action.

The plan records the overall intent and is revised when the model's understanding changes.

步骤 02

规划

第一次物理动作之前必须用 task_plan 提交简短的任务计划。

计划记录整体意图,模型的理解发生变化时随之修订。

STEP 03

Preview

robot_action first returns the planned trajectory and its cost without advancing the simulation.

The model checks the preview and adjusts the batch before anything moves.

步骤 03

预览

robot_action 先返回规划好的轨迹及其代价,不推进仿真。

模型检查预览结果,在任何运动发生前调整动作批次。

STEP 04

Execute

The cached plan is executed through the single gateway, which enforces permissions, state checks, budgets and stopping.

Simulation is paused while the model reasons and advances only through actions or explicit waits.

步骤 04

执行

缓存的规划经由唯一的执行网关执行,由网关统一负责权限、状态检查、预算和停止。

模型推理期间仿真暂停,只有动作或显式等待才会推进。

STEP 05

Feedback & evidence

After execution the harness returns images, measured robot state and an execution receipt.

Plans, commands, measured trajectories and scores are recorded separately; on stop the run is scored independently, archived and replayable.

步骤 05

反馈与证据

执行后运行框架返回图像、实测机器人状态和执行回执。

计划、指令、实测轨迹和评分分别记录;停止后独立评分、归档,并可回放。

Design principles设计原则

Four commitments四条原则

The model decides what to do, tools provide the capabilities, and the harness owns the boundary between decision and physical execution.

由模型决定做什么,工具提供能力,框架负责决策与物理执行之间的边界。

  1. 01

    Model-led, tool-assisted模型主导,工具辅助

    The model decomposes the task, picks tools and revises its strategy from feedback. Tools expose general capabilities and never ship a task-specific solution.

    模型负责拆解任务、选择工具,并根据反馈调整策略。工具只提供通用能力,不内置针对具体任务的解法。

  2. 02

    Closed loop on feedback基于反馈的闭环

    Candidate trajectories can be previewed before motion; afterwards the harness returns images and measured robot state, so the model can continue, look again or change plan.

    执行前可以预览候选轨迹;执行后框架返回图像和实测的机器人状态,模型据此决定继续、重新观察或修改计划。

  3. 03

    One gateway for every action所有动作经由同一网关

    Direct actions, model-written programs and optional policy models all execute through the same entry point, which enforces permissions, state checks, budgets and stopping.

    直接动作、模型编写的程序以及可选的策略模型都通过同一个入口执行,由该入口统一负责权限、状态检查、预算和停止。

  4. 04

    Traceable evidence可追溯的证据

    Plans, commands, measured trajectories and task scores are recorded separately. Motion finished, program exited and task succeeded stay distinct; failures and unknown states are kept.

    计划、指令、实测轨迹和任务得分分别记录。“运动完成”“程序退出”与“任务成功”三者分开记录;失败和未知状态同样保留。

Control paradigms控制范式

Three ways to put a model on a robot让模型控制机器人的三种方式

The paradigms differ in who decides, what reaches the actuators, and how failures are noticed and repaired.

三种范式的区别在于:由谁决策、什么指令到达执行器,以及如何发现和修复失败。

  (a)Action model (VLA)动作模型(VLA) (b)General model, direct control通用模型直接控制 (c) · ours本方法Auriga
Decides决策 A learned policy maps images and the instruction to actions.由学习得到的策略将图像和指令映射为动作。 A general model chooses the next action at each step (e.g. GPT Direct).通用模型在每一步选择下一个动作(如 GPT Direct)。 A persistent agent session with a task plan, tools and operation memory.一个持续运行的智能体会话,带有任务计划、工具和操作记忆。
Executes执行 Predicted action chunks go straight to the robot.预测的动作块直接下发给机器人。 One tool call, one robot step.一次工具调用对应机器人的一步动作。 Batched single- or dual-arm actions through one gateway: preview, then execute.单臂或双臂的批量动作经由同一网关执行:先预览,后执行。
Feedback反馈 Implicit, through the next camera frame.隐式反馈,通过下一帧相机图像获得。 A new observation after each call.每次调用后获得一次新的观察。 Images, measured joint / gripper state and execution receipts.图像、实测的关节 / 夹爪状态以及执行回执。
Recovery恢复 Bounded by what the training data covers.受限于训练数据的覆盖范围。 Re-decided at each call from the new frame; no plan preview or execution receipt.每次调用根据新图像重新决策;没有计划预览和执行回执。 Re-observe, inspect history, revise the plan; budgets and stopping enforced by the harness.重新观察、查看历史、修订计划;预算与停止由框架强制执行。
Demos演示

Selected cases, by capability按能力分组的精选 case

Six cases from the 50-case protocol, grouped by what each one tests. Pick a capability, watch Auriga's episode, then open the side-by-side comparison with a baseline on the same seed and layout.

从 50 case 协议中选出的 6 个 case,按考察的能力分组。选择一项能力,观看 Auriga 的运行,再打开与基线在同一随机种子与布局下的并排对比。

Compare with baseline与基线并排对比

case

Baseline基线
Auriga
Baseline基线 · GPT-as-Policy

Same case, same seed and layout.

同一 case,相同的随机种子与布局。

Tool layer工具层

General tools, structured results通用工具,结构化结果

Tools return structured results with images or file references attached as needed. The available set depends on configuration and observation settings.

工具返回结构化结果,并按需附带图像或文件引用。可用的工具集合取决于配置与观察设置。

Click a tool to see what it does, what goes in and out, and an example call.点击工具查看用途、输入输出与调用示例。

core tool核心工具optional, enabled by configuration可选,按配置开启

Getting started快速开始

Run one task运行一个任务

This starts a real model session and a robot simulation. Use a free GPU and a new run directory.

以下命令会启动真实的模型会话和机器人仿真。请使用空闲的 GPU 和新的运行目录。

mamba run -n Auriga python -u -m auriga.gateway \
  --benchmark --task general_pickup --seed 0 --layout-id 0 \
  --gpu 4 --port 8767 --run-dir runs/my-first-task
# open http://127.0.0.1:8767 for the workbench
# replay an existing run without a model or GPU: --replay --run-dir <run dir>
  • SimulatorRoboDojo with Isaac Sim installed, in a dedicated environment.
    仿真器在独立环境中安装 RoboDojo 与 Isaac Sim。
  • HardwareA CUDA GPU with free memory for the simulation.
    硬件一块有足够空闲显存运行仿真的 CUDA GPU。
  • Model accessThe native Codex runtime, with authentication and model access configured.
    模型访问原生 Codex 运行时,并已配置认证与模型访问权限。
Citation文献引用

Cite this work引用

Until the paper is out, please cite the project page.

论文发布前,请引用本项目页。

@misc{auriga2026,
  title        = {Auriga: A Model-Led Agent Harness for Embodied Manipulation},
  author       = {StellarEdge Team},
  year         = {2026},
  howpublished = {Project page},
  note         = {StellarEdge Team}
}