Robotics data, explained.
Before you list, it helps to know what a lab is looking at. Here's what buyers actually train on, the specs they check first, and how to describe what you've collected so they can evaluate it.
Two guides, depending on where you are
This page is the reference: the specs, the vocabulary, and what a buyer checks. If you are earlier than that and still working out whether your work is worth recording at all, start with one of these.
How to sell your trade data
Which trades labs want, what raises the price of a trade dataset, the rights you need before you record on a customer's site, and the seven steps from a work day to a live listing.
How to sell trade data →How to earn from your data
Which data actually pays and which pays nothing, what an hour is worth once the middleman is gone, and the four routes to earning from what you already record.
How to earn from your data →The Butter Kit
A wearable SLAM tracking device that records pose and ego-motion alongside video, which is the single largest lever on what an hour of capture earns.
See the device →What labs are actually training
Robotics labs can generate almost anything in simulation, and they still can't fabricate the real world: real lighting, real clutter, real hands doing skilled work, real mistakes being corrected. That's the gap your data fills, and it's why three fairly different kinds of buyers want three fairly different things.
Robots that handle things
Arms, grippers, and humanoids learning physical tasks like wiping a counter, sorting parts, or loading a machine. They want first-person video of skilled hands doing real work, with the motion that produced it.
Looks forFirst-person RGB, many short episodes, task labels, hands in frame.
Robots that move through space
Delivery robots, inspection drones, warehouse movers, autonomous equipment. They need to know how a space is laid out and how a body actually travels through it, including around people and clutter.
Looks forContinuous takes, pose/trajectory, depth or LiDAR, varied sites.
Systems that predict what happens next
Perception stacks, 3D reconstruction, and simulators trained to model real environments. They put a premium on long, uninterrupted, spatially grounded footage of ordinary places.
Looks forLong sessions, dense geometry, consistent calibration, no cuts.
What a buyer checks, roughly in order
Every dataset gets skimmed against the same short list. You don't need to hit every item on it. State honestly where you land, so a buyer can tell in thirty seconds whether your data fits their model.
- HoursTotal recorded duration. Every buyer sorts on it first, though it rarely tells the whole story on its own.
- EpisodesHow many discrete takes those hours are split into. 200 hours across 600 episodes is worth far more than 200 hours in 4 sittings: more starts, more variation, more attempts at the same task.
- ModalityWhich sensor streams are actually in the file: RGB, depth, LiDAR, IMU, GPS, force-torque. List what you captured, even when your device can record more.
- Resolution & rateFor example 1080p at 30 fps video, LiDAR at 10–20 Hz, IMU at 100–200 Hz. Buyers mostly want the numbers stated plainly and held consistent across the whole set.
- Pose / SLAMWhether every frame carries a position and orientation in the scene. Pose lets a buyer place each shot in the room, and they price for it.
- Time syncAll streams on one clock, with a timestamp per frame. Unsynchronized streams are the most common reason a technically interesting dataset gets passed over, because the buyer can't line the video up with the motion.
- CalibrationCamera intrinsics, plus the fixed transform between sensors (extrinsics). Without it, depth and LiDAR points can't be projected onto the image reliably.
- DiversityHow many distinct sites, lighting conditions, times of day, and operators. Ten kitchens beat one kitchen ten times over, at the same hour count.
- Task labelsWhat is happening, and when. Even coarse start/stop markers per task, like “0:00–4:12 loading dishwasher,” move a dataset up a tier, because the buyer doesn't have to pay someone to segment it.
- FailuresKeep the fumbles, drops, and re-tries. Models learn correction from them, and a set of only clean successes teaches a robot nothing about recovering.
- FormatConsistent file layout and naming, a manifest of what's in each episode, and a short README. MP4 plus CSV/JSON, ROS bags, and MCAP all work fine. Inconsistency is what hurts.
- RightsYou have the right to license the footage, faces and plates and screens are handled, and no customer PII rides along. Buyers ask about this before they ask about anything else on this list.
The words buyers use
You'll see these in buyer requests and in our listing form. None of them are as complicated as they sound.
- Episode / trajectory
- One continuous take, from record to stop. Manipulation buyers usually count episodes before hours.
- Demonstration
- An episode where a person does the task correctly on purpose, so a policy can learn to imitate it.
- Ego-motion
- The movement of the camera itself. Its path through the scene is what makes first-person data trainable.
- Egocentric
- Shot from the first-person point of view, with the camera worn on the body while the work happens.
- 6-DoF pose
- Position (x, y, z) plus orientation (roll, pitch, yaw). Six numbers that say exactly where a sensor was and which way it faced.
- Point cloud
- A set of 3D points, typically from LiDAR or depth. It carries geometry, with no color or labels attached.
- IMU
- Inertial measurement unit: accelerometer plus gyroscope, sampled fast. It fills in motion between camera frames.
- Odometry
- Estimated motion from one moment to the next. Accurate step to step, but drifts as errors accumulate.
- Drift / loop closure
- Drift is a tracked path slowly bending away from truth. Loop closure is the correction when the system recognizes a place it's already been.
- Intrinsics / extrinsics
- Calibration. Intrinsics describe one camera's lens; extrinsics describe how sensors sit relative to each other.
- Force-torque
- Measured push and twist at a contact point. Rare in wearable capture, prized for contact-rich tasks like inserting or wiping.
- Proprioception
- A robot's own joint positions and velocities. It shows up in data a robot records about itself, so wearable capture rarely includes it.
- Teleoperation
- A human driving a real robot to produce training data. It costs a lot per hour, which is one reason labs look to wearable capture.
- VLA model
- Vision-language-action: a model that takes in pixels and an instruction and outputs robot motion. The main consumer of this market's data.
- Ground truth
- The measurement treated as correct, used to score everything else, such as a surveyed position or a hand-checked label.
- Sim2real
- Training in simulation and deploying on hardware. Real data is what closes the gap simulation leaves behind.
Same hours, very different value
Labs pay roughly $7–$15 an hour for expert data today, and that band covers a lot of ground. Where you land inside it, or above it, comes down to how structured, how varied, and how usable your set is the day it arrives.
- Many episodes of the same task, done by different people, in different places.
- Pose or SLAM tracking recorded alongside the video.
- Timestamps on every stream, from one clock.
- Task segments marked, even roughly.
- Rare or hard-to-stage work: trades, industrial settings, specialist procedures.
- Mistakes and recoveries left in.
- Clean rights: your premises, your staff, consent handled.
- Hours of a single static viewpoint with almost nothing happening.
- Streams that can't be aligned in time.
- Modalities claimed in the listing that aren't in the files.
- Inconsistent naming, missing episodes, no manifest.
- Heavy compression that smears the detail a model needs.
- Faces, screens, documents, or customer data you can't license.
- One long unbroken take with no structure and no labels.
Put the specs in the title
Buyers scan a marketplace the way you'd scan a parts catalog. The listing form asks for hours, episodes, modalities, SLAM, robot platform, environment tags, and tasks. Fill in every one you can, then repeat the important ones in the title and summary.
“Restaurant kitchen footage”
Nothing here tells a buyer whether it fits their model. How long, how many takes, worn or mounted, tracked or not?
“Commercial kitchen prep: 120h RGB-D + SLAM, 340 episodes, 3 sites”
Task, setting, hours, modalities, pose, episode count, and diversity, all before the buyer clicks.
- What work is being done, and by whom: the trade, or the daily routine.
- Where it was shot, how many distinct places, and over what period.
- Exactly which streams are in the files, and at what resolution and rate.
- How it's organized: episode structure, file format, any labels.
- What's been done about privacy, covering faces, screens, and customer information.
Ready to list?
Set your own price, submit for review, and sell direct to frontier and neo labs, with no aggregator in the middle.
Already collecting? Go to your data portal
Need a device first? See the Butter Kit