autonomylabs.dev  /  field log  /  001
Field Log · Entry 001
First walks: the brain moves before the wheels do
21 Sep 2026
Peoria, Illinois
Indoor · two runs
BrainQwen3-VL 2B · 4-bit · on-device ComputeJetson Orin Nano 8 GB · 25 W Eyes4 × USB cameras at 90° Cloudnone
Summary

Photons in, decisions out. Motion still mocked.

We strapped four cameras and a Jetson to a hand cart and pushed it through a house, twice. The on-board vision-language model made a decision every 1.5 seconds, sustained, with every reply parsed. It noticed a person every single time one was near. It also mistook the cart's own handle for an obstacle, could not tell front from back, and, told the mission in one sentence, went from backing up 62% of the time to never backing up at all, including when its nose was against the kitchen cabinets.

The finding that matters: the model's perception is good and improves with context. Its policy swings with the wording of a prompt. So the model will describe the world, and a small set of testable rules will decide what to do about it.

The rig

Four webcams on a drum, a USB hub and a Jetson Orin Nano strapped to a blue hand cart, one extension cord running off it.
The whole autonomy stack, minus wheels. Cameras on a drum at 90°, hub, Jetson, one extension cord. A human provides the motion.
Four camera views side by side: front door and yard, a blank wall, the cart handle, a living room.
What the four cameras see, front, right, rear, left. They capture within a millisecond of each other and are tiled into one 2×2 image per decision.

Two runs, one sentence apart

Run 1 · no missionRun 2 · one sentence of mission
Decisions85 in 140 s59 in 104 s
Per decision (median)1.57 s1.76 s
proceed / hold / reverse / stop / turn0 / 24 / 53 / 5 / 359 / 0 / 0 / 0 / 0
Garbled replies00 (7 verbose, since parsed)

Each cell below is one decision, left to right in time. Hover for what the model said.

proceedholdreversestopturn right
Run 1 · "choose a safe action"
Run 2 · same route, plus: "we are moving forward to the kitchen; the person at the rear handle is the operator; only react to things close ahead; do not reverse"

What the model saw

2x2 camera mosaic: empty hallway ahead, cart handle in the rear view.
Run 1, frame 0. Empty hallway. Verdict: reverse, "obstacle in the rear view". The obstacle is the cart's own handle.
2x2 camera mosaic: a person holding the cart handle in the rear view.
Run 1, frame 14. Person at the handle. Verdict: hold, confidence 0.95. Correct, and repeated for most of the walk.
2x2 camera mosaic: kitchen cabinets close in front, nobody in frame.
Run 1, frame 76. Parked at the cabinets, nobody around. Verdict: reverse, 0.8, twenty frames running. Nothing changed; the verdict never did either.
2x2 camera mosaic: hallway ahead, person at the handle behind.
Run 2, frame 12. Same person, same handle. Verdict: proceed, "person at rear handle". One sentence turned a hazard into the operator.
2x2 camera mosaic: cabinet doors inches from the front camera.
Run 2, frame 57. Nose to the cabinets. Verdict: proceed, 0.95, "clear view of kitchen". Same brain as frame 76 above, one sentence apart, opposite mistake.

What it settles

Next

Stack for the record: Ollama 0.13.3 with flash attention, Qwen3-VL 2B Instruct Q4_K_M, a four-camera terse-mosaic harness, one decision per ~30-token reply. Every frame and every decision is archived per run.

Autonomy Labs · Field Log 001Home  ·  Architecture  ·  Training walk  ·  Build with us