BrainQwen3-VL 2B · 4-bit · on-deviceComputeJetson Orin Nano 8 GB · 25 WEyes4 × USB cameras at 90°Cloudnone
Summary
Photons in, decisions out. Motion still mocked.
We strapped four cameras and a Jetson to a hand cart and pushed it through a house, twice. The on-board vision-language model made a decision every 1.5 seconds, sustained, with every reply parsed. It noticed a person every single time one was near. It also mistook the cart's own handle for an obstacle, could not tell front from back, and, told the mission in one sentence, went from backing up 62% of the time to never backing up at all, including when its nose was against the kitchen cabinets.
The finding that matters: the model's perception is good and improves with context. Its policy swings with the wording of a prompt. So the model will describe the world, and a small set of testable rules will decide what to do about it.
The rig
The whole autonomy stack, minus wheels. Cameras on a drum at 90°, hub, Jetson, one extension cord. A human provides the motion.What the four cameras see, front, right, rear, left. They capture within a millisecond of each other and are tiled into one 2×2 image per decision.
Two runs, one sentence apart
Run 1 · no mission
Run 2 · one sentence of mission
Decisions
85 in 140 s
59 in 104 s
Per decision (median)
1.57 s
1.76 s
proceed / hold / reverse / stop / turn
0 / 24 / 53 / 5 / 3
59 / 0 / 0 / 0 / 0
Garbled replies
0
0 (7 verbose, since parsed)
Each cell below is one decision, left to right in time. Hover for what the model said.
proceedholdreversestopturn right
Run 1 · "choose a safe action"
Run 2 · same route, plus: "we are moving forward to the kitchen; the person at the rear handle is the operator; only react to things close ahead; do not reverse"
What the model saw
Run 1, frame 0. Empty hallway. Verdict: reverse, "obstacle in the rear view". The obstacle is the cart's own handle.Run 1, frame 14. Person at the handle. Verdict: hold, confidence 0.95. Correct, and repeated for most of the walk.Run 1, frame 76. Parked at the cabinets, nobody around. Verdict: reverse, 0.8, twenty frames running. Nothing changed; the verdict never did either.Run 2, frame 12. Same person, same handle. Verdict: proceed, "person at rear handle". One sentence turned a hazard into the operator.Run 2, frame 57. Nose to the cabinets. Verdict: proceed, 0.95, "clear view of kitchen". Same brain as frame 76 above, one sentence apart, opposite mistake.
What it settles
The loop is real. Capture, tile, reason, decide, log: 1.5 seconds, all on the board, no cloud. The plumbing is no longer the question.
Perception is the asset. People are noticed every time, in any camera. Context makes it better: told who the operator is, it stopped fearing him; in one pass it correctly waved through a person far down the hall.
Policy is not. Without a mission the model backs away from everything. With one it drives at everything. The prompt, not the scene, was choosing the action.
So: split the jobs. The model reports what is in each camera and roughly how near it is. Plain rules turn that into motion. A person near ahead means hold, a wall near ahead means stop, a clear path with a mission means go. Rules are testable. Prompts are moods.
Field rigs fail at the plug. One attempt ended at 24 frames when the extension cord wiggled out. Tape, strain relief and a loop of slack are now on the checklist.
Next
Perception-only output plus rules, on the same hallway route, against these two traces.
Rolling memory on, so each decision knows what the last few frames said.
The yard: trees, a road edge, a distant house, weather. Then wheels.
Stack for the record: Ollama 0.13.3 with flash attention, Qwen3-VL 2B Instruct Q4_K_M, a four-camera terse-mosaic harness, one decision per ~30-token reply. Every frame and every decision is archived per run.