N NavHarness

NavHarness
Adaptive Goals for
Vision-and-Language Navigation

79.0%R2R-CE success rate+5.0 points over Codex CLI
82.5%RxR-CE success rate+6.2 points over Codex CLI
83.3%Real-world success rate8 routes · 24 trials · 1.51 m NE
METHOD PERFORMANCE

Navigation performance

R2R-CE and RxR-CE results. NE is in meters; the remaining metrics are percentages.

M = monocular, P = panorama, D = depth; 3-cam = three fixed RGB cameras. Results retain their original protocols; SmartWay and AgenticNav use Open-Nav’s 100 episodes. Bold marks the best reported value per metric; underlining marks the best baseline within each group unless bold. Shaded rows denote NavHarness without memory compression. Dashes denote unreported or pending results.

01 / ABSTRACT

Keep local actions aligned
with the entire route.

Vision-and-Language Navigation requires embodied agents to generate actions from instructions and observations. General-purpose multimodal agents provide a promising basis, but plausible local actions do not ensure that execution remains consistent with the intended route, particularly over long horizons. Accumulated interaction history also increases the input required for later decisions.

NavHarness organizes navigation around adaptive goals. A Goal Agent sets local objectives, a Visuomotor Agent executes them, a Verify Agent checks observable completion conditions, and a Memory Agent compresses completed multimodal history while retaining information needed for subsequent navigation.

We evaluate R2R-CE and RxR-CE, compare variants across three model backbones, and study context evolution. In real-world evaluation, NavHarness achieves 83.3% success and 1.51 m navigation error across eight challenging routes evaluated three times each.

02 / METHOD

Four agents.
One goal-grounded navigation loop.

NavHarness framework: adaptive goal generation, visuomotor execution, completion verification, and progress-aware memory
The NavHarness framework.
01

Goal Agent

Set a local goal and an observable completion question from the instruction, current view, and history. Revise the goal when the scene requires it.

02

Visuomotor Agent

Translate the active goal into forward movements, turns, and stopping, using current observations and verification feedback.

03

Verify Agent

Check observed outcomes against the goal’s completion condition. Guide continued execution, correction, or advancement.

04

Memory Agent

At verified completion, consolidate the finished stage into a compact route summary, verification outcomes, and visual keyframes.

03 / SIMULATION RESULTS

Consistent gains across three backbones.

On R2R-CE, the framework improves success rate with Qwen-3.6-plus, GPT-5.6-Sol, and GPT-6-Astra. Compression retains part of the gain.

R2R-CE success rate across three models, comparing Codex CLI, NavHarness, and NavHarness with compression
04 / PROGRESS-AWARE MEMORY

Compress the completed past.
Keep the active goal in view.

Verified goal completion provides recurring boundaries for summarizing multimodal history.

Text per decision falls by 55.5–71.6% across the four benchmark/backbone settings. Total-token savings depend on the backbone and come with navigation trade-offs.

GPT-5.6-Sol · RxR-CE56.59%

fewer total tokens per episode

With goal-completion compression compared with the same harness without compression.

Success rate34.0 → 32.0%SPL25.0 → 21.9%

For GPT-6-Astra on RxR-CE, total tokens decrease from 1538.84k to 1476.45k (4.05% savings), with SR 82.5 → 81.0% and SPL 64.7 → 58.3%. Total tokens include internal calls and compression overhead; incomplete usage records are excluded.

COMPLETE EFFICIENCY RESULTS

Navigation quality and context consumption

Cache, Input, and Total are thousands of tokens per episode. Image/Dec. is the mean number of input images per external decision; Text/Dec. is the mean input text length in thousands of characters per external decision. Total ratio is Goal/Off (%).

ModelHarnessCompress.SR ↑SPL ↑CacheInput ↓Total ↓Image/Dec. ↓Text/Dec. (k) ↓Total ratio (%) ↓
R2R-CE
GPT-5.6-SolCodex CLINative62.644.51351.401398.721410.3112.4368.99—
NavHarnessOff67.047.91234.501334.641345.2313.1080.57—
NavHarnessGoal66.050.3770.45851.85861.8211.8032.3164.06%
GPT-6-AstraCodex CLINative74.062.5656.30687.77689.819.4273.72—
NavHarnessOff79.066.6510.55577.91580.9111.2881.15—
NavHarnessGoal77.065.6478.34544.82548.877.2136.1294.48%
RxR-CE
GPT-5.6-SolCodex CLINative28.021.01946.312017.162027.8134.5978.65—
NavHarnessOff34.025.04008.334277.804298.5545.05122.09—
NavHarnessGoal32.021.91702.701847.591865.9218.1434.6543.41%
GPT-6-AstraCodex CLINative76.360.01097.851145.121149.5026.6781.61—
NavHarnessOff82.564.71392.891533.901538.8427.26105.92—
NavHarnessGoal81.058.31326.071467.081476.4513.6238.5095.95%

Cache is included in Input. Total additionally includes output. Off disables compression; Goal uses verified goal boundaries; Native denotes baseline context management. Total ratio compares Goal with Off; a smaller ratio indicates fewer total tokens.

05 / REAL-WORLD NAVIGATION

Follow the instruction.
See the entire recorded route.

Indoor and outdoor navigation on a Unitree Go2, with RGB observations from an Intel RealSense D435i.

Fixed-rate observation sequence · 2 fpsDownload MP4 ↓
NAVIGATION INSTRUCTION

The yellow chair with a white ball

Explore the routes

8 cases
PHYSICAL EVALUATION

8 routes, each tested 3 times.

Approximately 15–20 m per route, across corridors, laboratories, classrooms, sofa areas, stairs, and outdoor spaces.

Real-world evaluation · manuscript Table 3
MethodSR ↑NE (m) ↓
NaVILA20.83%3.91
AwareVLN20.83%3.89
StreamVLN25.00%3.78
JanusVLN50.00%2.34
NavHarness83.30%1.51
06 / SIMULATION DEMOS

See NavHarness navigate
from the agent’s point of view.

Eight successful M01–M12 trajectories selected from the recorded R2R-CE and RxR-CE evaluations for low navigation error and strong path efficiency. Every clip uses saved first-person observations from the original run; no navigation episode was rerun.

Loading trajectories