AGI for robots that learn & improve on-edge.
Two systems, one robot: an on-board governor that makes your robot faster and better by itself after it ships — and RL environments that pre-train it to cross the reality gap on the first try.
Measured against current RL pipelines
Cost-effective. GPU-efficient. Self-improving.
Time to fine-tune · vs. training data
The more data you train on, the further Cadenza pulls ahead, fine-tuning in a fraction of the time current RL systems need.
01 · The systems
Two systems. One self-improving robot.
Pre-train a policy on RL environments that actually cross the reality gap — then let it keep getting better on the robot itself, on-edge, long after it ships.
Robots that scale & learn by themselves
An on-board governor watches the robot succeed and makes it faster and better — by itself, after it ships. No retraining, no cloud.
- Self-improves in place. Tiny, committed speed-ups per task while the robot runs.
- Safety-gated. Keeps a change only if the outcome holds; reverts the instant it wouldn't.
- Long-horizon. Microstep checkpoints turn long tasks into sub-goals it solves one by one.
RL environments that cross the gap
Novel, GPU-efficient RL environments that auto-calibrate physics against real telemetry, so policies are deploy-ready before they touch hardware.
- Sim-to-real fidelity. Domain randomization tuned to your captured robot telemetry.
- Linear scaling. 100k deterministic instances per GPU partition, one seed.
- Any robot. URDF, MJCF, or a Cadenza spec — arm, hand, or quadruped.
02 · Pre-training
Pre-train policies on RL environments in a dozen lines.
Name the robot and the scene, and Cadenza builds the RL environment, auto-calibrates physics against real telemetry, and scales to hundreds of thousands of parallel instances. The policy you pre-train here is the same arm that keeps improving itself on-edge below.
- Declarative tasks. Reward functions defined inline and versioned, not glued code.
- Sim-to-real fidelity. Domain randomization tuned to your captured robot telemetry.
- Linear scaling. 100k deterministic instances per GPU partition, one seed.
1import cadenza as cz2 3# Build an RL environment for the Cadenza arm on Cadenza Lab.4# Physics auto-calibrates against captured telemetry.5env = cz.Environment(6 robot=cz.robots.Arm(dof=6),7 backend="mujoco",8 scene="lab/stacked_blocks",9 fidelity="sim2real", # domain-randomized, contact-rich10)11 12@env.task("pick_place")13def reward(state, action):14 placed = cz.metrics.object_at(state, target="left")15 stable = cz.metrics.tower_intact(state)16 effort = cz.metrics.energy(action)17 return 4.0 * placed + 2.0 * stable - 0.01 * effort18 19# Scale to 100k parallel instances on one GPU partition.20fleet = env.parallelize(n=100_000, seed=7)21 22policy = cz.train(23 fleet,24 algo=cz.algos.PPO(clip=0.2, gae_lambda=0.95),25 steps=2_000_000,26)27 28policy.export("arm_pick_place.cadenza") # ships with the on-edge governor03 · On the edge
Self-improvement that ships with the robot.
Give the arm a goal in plain English. It resolves the command into the action library, tokenizes it into microstep checkpoints, and runs the task on a delicate three-block tower. Then the on-board governor makes it faster on its own, rep over rep — keeping every speed-up that leaves the tower standing, backing off the instant one wouldn't. No retraining, no cloud.
SemanticLayer
natural language → action library
MicrostepTokens
per-phase checkpoints to solve
On-edge governor
speeds up; reverts if the tower shifts
1import robogpt2 3arm = robogpt.connect("cadenza-arm") # 6-axis, on-edge governor4plan = arm.understand(PROMPT) # NL → action library5plan = robogpt.microsteps(plan) # per-phase checkpoints6 7for ep in arm.run_forever(plan): # self-improves in place8 if ep.clean and ep.faster:9 arm.commit(ep.speeds) # keep the speed-up10 else:11 arm.revert() # tower shifted → back off04 · Who it's for
Built for teams shipping self-improving robots.
For inference startups
A rebuilt rollout engine reclaims 70% more GPU headroom, so you can serve more robot policies per partition without buying more silicon.
For robotics developers
Everything to take a policy from sim to a shipping robot: RL environments, data, evals, and an on-edge governor that keeps it improving — in one layer.
GPU-efficient by design
Cost-effective rollouts that scale linearly to 100k deterministic instances per GPU partition, one seed.
Reliable across the gap
Contact-rich physics auto-calibrated against captured telemetry, so policies cross the reality gap on the first deploy.
Any robot description
URDF, MJCF, or a Cadenza robot spec. Bring an arm, a hand, or a quadruped, and the systems adapt to your hardware.
Self-improving in the field
Export one artifact that runs in sim and on hardware, with the on-edge governor making it faster and better after it ships.
Ship robots that
improve themselves.
Two systems for self-improving robots: RL environments to pre-train them, and an on-edge governor that keeps them getting better in the field. Request access and we'll stand it up with you.