A scene can look plausible and still be spatially wrong.
Agentic reconstruction uses a large model to inspect observations, call modeling tools, and iteratively build and check an editable scene. The challenge is not only to produce a convincing room, but to preserve the scale, shape, and spatial relationships that make it useful beyond rendering. AWSM asks how geometric evidence can constrain this process so that its outputs become more faithful spatial references for downstream agents.
We reconstruct World Lobby, a simulated scene in NVIDIA Isaac Sim, through four complete engineering routes, from RGB-only modeling to ground-truth-pose-conditioned modeling. Each route produces a frozen Blender scene, not merely a point cloud or a novel-view renderer. The experiment asks what additional geometric evidence buys, where it fails to help, and which errors survive object-centric reconstruction.
From plausible reconstruction to a usable spatial reference
A reconstruction agent can assemble a plausible room while getting a corner, distance, or passage wrong. For an embodied agent, those are not merely visual defects: they change where a destination lies and which route a body can follow. Geometry grounding means constraining the modeling process with pose, depth, and metric evidence where available, rather than relying on visual plausibility alone. The question is how those constraints survive the conversion from observations into editable objects.
The reception-desk task makes this connection concrete: mark a destination in the reconstructed map, specify routes, and execute them with robot controllers. We use ‘scene representation’ for the editable objects and geometry, ‘map’ for their navigation role, and ‘simulation assets’ for their integration into a simulator. The ‘world’ in AWSM refers to this reusable spatial environment, not a learned dynamics predictor. A persistent spatial foundation connecting navigation, memory, interaction, and simulation is the longer-term goal; the experiments and demo examine its geometric basis and one downstream use.
Contributions
A four-route study of geometry-grounded agentic reconstruction, keeping the scene, 180 RGB identities, modeling objective, and editable output format fixed while comparing complete pipelines with different geometric evidence.
An evaluation protocol that separates trajectory error, native depth, final-scene depth, surface geometry, and rendered appearance, with frozen Blend/GLB assets, reproducible evaluation subsets, public hashes, and interactive comparisons.
A map-based multi-robot simulation example connecting a reconstructed scene to destination marking, predefined routes, and controller execution. It illustrates downstream use rather than a controlled navigation comparison across reconstruction routes.
Geometry grounding: scale is necessary, but not sufficient
Monocular projection geometry leaves a global scale ambiguity; learned metric depth adds a useful prior, not an independent measurement of every real-world dimension. Calibrated RGB-D, a known baseline or length, and visual-inertial estimation can supply metric constraints. IMU fusion can recover metric motion under sufficient excitation and reliable initialization, bias estimation, synchronization, and calibration. An IMU is not a direct distance sensor, and its presence alone does not guarantee accurate scale.
Scale, local shape, camera pose, and navigable connectivity must be checked separately. A curved corner simplified into a square cannot be repaired by rescaling alone. Pose-conditioned depth can improve reconstruction, but neither the DA3 ablation nor our M2/M3 results imply that every depth metric improves monotonically. The four routes below are system comparisons, not an isolated IMU ablation.
The reconstruction agent follows a shared observe–build–verify loop: inspect the available evidence, write Blender Python to construct objects, render review views, and revise the scene before freezing it for evaluation. This tool-using modeling process is what ‘agentic’ describes here; the downstream robot controllers have a separate role.
All routes receive the same 180 RGB identities and the same Astra–Blender modeling objective. M1 tests how far visual priors can go without metric measurements. M2 adds RGB-only ViPE poses and pose-conditioned DA3 depth. M3 uses monocular-inertial ORB-SLAM3 poses with the same class of DA3 observations. M4 supplies ground-truth camera poses to diagnose the remaining modeling error; it does not receive the GT mesh or GT depth as modeling input.
Table 1.
Agentic reconstruction loop
01 · ObserveInspect all 180 RGB frames and, where allowed, pose-conditioned DA3 geometry.
02 · BuildWrite Blender Python that decomposes the room into named objects, materials, cameras, and collision proxies.
03 · VerifyCompare paired RGB and optical-Z renders at fixed review cameras; revise within a fixed version budget.
04 · FreezeExport Blend/GLB, manifests, semantic IDs, and hashes before any GT score is visible to the author.
05 · EvaluateMeasure pose, native depth, model depth, surface distance, and appearance under one frozen protocol.
Ground truth is introduced only after the scenes are frozen. M2 and M3 are evaluated under one global SE(3) alignment with scale fixed to one; no mesh ICP or per-view fitting is used. M1 has no native metric trajectory, so its model-space numbers use one frozen GT-assisted Sim(3) from five manual camera associations and remain diagnostic rather than evidence of metric recovery.
04 / RESULTS AND ANALYSIS
Three findings matter more than a leaderboard
FINDING 01
1. Better pose does not mechanically imply better native depth
M3 improves trajectory accuracy over M2—ATE falls from 0.167 m to 0.121 m and rotation error from 0.31° to 0.19°. Yet its native DA3 AbsRel is slightly worse (20.38% versus 19.06%). After modeling, the ordering reverses: M3 reaches 8.37% model-depth AbsRel versus 9.29% for M2. Upstream metrics and downstream scene fidelity are related, but they are not interchangeable.
Table 2.Figure 1. The upstream trajectory and native-depth ranking does not map one-to-one onto frozen-scene depth.
FINDING 02
2. Geometric evidence turns a plausible layout into a more faithful space
Under the declared alignment and without mesh ICP, bidirectional surface error falls from 0.376 m for the RGB-only diagnostic baseline to 0.118 m with ViPE, 0.082 m with ORB-SLAM3, and 0.072 m with GT poses. M4 is strongest on geometry in this case, while M3 closes most of the gap using an estimated trajectory.
Table 3.
FINDING 03
3. Qualitative comparison
Inspect the reconstructed scenes, not just aggregate scores
Use the shared camera to compare each reconstruction with GT. Start with M1 to see how a semantically plausible room can drift in metric layout; compare M2 and M3 to find regions where geometry improves without a matching gain in RGB similarity; then inspect M4 to isolate errors that remain even when camera pose is no longer the bottleneck.
Figure 2. M1–M4 and input GT across five fixed views.
Rotate, zoom, and inspect the frozen scenes
M4GT
Scroll here to load the 3D comparison.
Figure 3. One camera and one full viewport compare the selected frozen model with GT. Ceilings and surrounding walls are hidden at runtime for the same cutaway-style inspection; source files remain unchanged.
M1 / VIEW 0
INPUT / GT RGB
Figure 4. Frozen-model rendering and input GT RGB from exactly the same view.
The browser hides ceilings and surrounding walls only for cutaway inspection. Runtime registration and display choices do not modify the downloadable frozen GLB or Blend files.
What geometric grounding changes—and what it does not
Modeling behaves like a structured—and lossy—regularizer
For M2–M4, the final scene depth is substantially closer to GT than native DA3 depth on the shared 180-view protocol: AbsRel falls by roughly 51–60%. One plausible interpretation is that object-level aggregation suppresses inconsistent local predictions. This is a hypothesis, not a controlled causal result: the same abstraction can also erase real surfaces and fine geometry.
Measurement constrains generation; it does not solve appearance
The strongest geometric route does not dominate PSNR or SSIM. Pose removes one source of uncertainty, but appearance still depends on material estimation, lighting, object detail, and renderer mismatch. Evaluating only RGB would miss spatial progress; evaluating only geometry would miss perceptual failure.
The 100-view check measures consistency, not unseen-view generalization
Table 7 uses a seeded set of 100 cameras drawn from the 175 modeling frames outside the fixed-five check. These are additional evaluation views relative to that check, but their RGB images were available during modeling. On this set, M4 has the lowest depth AbsRel (7.47%), while M3 has the lowest RMSE (1.090 m).
Figure 5. Model-depth error at the fixed-five check; teal marks missing or invalid predictions.Table 6.Table 7.
Beyond the lobby: what survives in a phone-captured space?
Office Café is not a second run of the World Lobby benchmark and is not quantitative validation under the same protocol. It is an external, real-capture case study showing how an object-centric scene, a solved source-camera trajectory, video-to-model comparison, and physics replays can coexist in one inspectable artifact. An author-observed failure—a curved real-world corner simplified into a square one—motivates separating local shape from global scale. A correct scale alone cannot repair that shape error; this is a qualitative observation, not another measured benchmark.
Local preview from the published model-render video.
Shake experiment
Local preview from the published strong-shake replay.
Links request English with lang=en; the external host controls language and availability. These local content previews are not screenshots of the external website. If embedding is unavailable, use the full-screen links or the local models and videos below.
Current modelVideo plus estimated depth as input; drag to orbit, right-drag to pan, and scroll to zoom.Legacy modelVideo-only input; this legacy layout was not recalibrated against the ViPE scan.
Both models open in an elevated cutaway view: upper geometry is clipped only for display to expose the interior. Toggle Cutaway to inspect the complete model; Reset view restores the overview. Authored coordinates, materials, and model files are unchanged. Linked orbit controls compare layouts, not registered geometry.
Input videoOriginal phone walkthrough.Procedural modelModel rendered along the source-camera trajectory.Point cloudScan rendered along the same trajectory.
Shared playback · 70.1 seconds · 30 fps
Strong earthquake effectEvery object in the scene is interactable, and cabinet doors can open; this replay shows the physical response under strong shaking.
DEMO / MAP-BASED EMBODIED EXECUTION
A reconstructed map. A destination. Robots in motion.
The demo connects reconstruction to downstream use in NVIDIA Isaac Sim. The reception desk is marked in the navigation map, routes are specified, and a drone, humanoid, quadruped, and wheeled robot execute the task. The same spatial reference connects a destination, route constraints, and motion control: a scene to inspect becomes a map to act with.
Task intent: Send the robots to the reception desk and have them line up there.
01Task instruction
02Locate the reception desk
03Set routes on the map
04Execute with controllers
Map-based navigation in simulation · predefined routes · simulator-pose feedback. The authors confirm that the demonstration has been manually reviewed.
Navigation map. The reception desk and planned routes share the reconstructed scene. This is a static design preview, not a recorded trajectory. Open the image to inspect the full-resolution map.Demo protocol and scope
The instruction above describes the intended task, not a demonstrated language-to-plan parser. Routes were manually specified and screened against both the M4 reconstructed map and original geometry. This is not a reconstruction-only planning benchmark.
The 156-second presentation edit plays at 1×: Scene (0–6s), Reconstruction (6–12s), Planned route (12–18s), and Navigation (18–156s). Navigation uses the new run 20260930-214924; the opening 12.32 seconds of waiting are omitted, and the later terminal hold is outside this cut. The map is a static route design, not a measured execution trace.
This demo does not evaluate online visual localization, learned instruction understanding, memory retrieval, or real-robot transfer. Those are separate capabilities; the map-based execution shown here remains the intended demonstration.
Manual presentation review is separate from the historical automated checks retained below; it does not revise their recorded verdicts.
AWSM connects two roles: an agent that constructs an editable scene, and downstream agents that use it as a spatial reference. SLAM and visual-inertial estimation supply pose and metric constraints; ViPE and Depth Anything 3 contribute geometric observations. Concurrent work AHa-3D explores video-driven, tool-based Real2Sim, including camera and 3D-reference estimation from video. Our emphasis is geometric grounding: bringing additional physical measurements, especially IMU-informed visual-inertial pose and metric constraints, into agentic scene construction rather than relying on video-derived evidence alone. The four-route comparison tests reconstruction systems, not the isolated effect of IMU; the demo illustrates a map-based application, not a complete autonomous embodied system.
Our evaluation therefore spans both sides of the interface: upstream systems are judged as measurement pipelines, while the Blender outputs are judged as persistent spatial artifacts. This differs from evaluating pose, depth, or novel-view synthesis in isolation.
The current outputs are editable scene snapshots, and the demo shows a scene-specific simulation integration. The longer-term goal is to make geometry-grounded agentic reconstruction a reusable process for building and maintaining spaces that phygital agents can use. Saving an editable asset is a starting point, not a demonstration of lifelong map maintenance. The directions below extend the present evidence.
Navigate
Use destinations, free space, and body-specific clearance in a shared geometric frame. Better reconstruction can reduce map mismatch; reliable navigation also needs localization and current observations.
Remember & retrieve
Link object identities and locations to source views so language, images, and past observations can refer to the same place. Retrieval and safe arrival should be evaluated separately.
Interact & maintain
Edit object properties and relationships, revisit changed areas, and preserve evidence and versions. Persistent memory must distinguish observed geometry, inferred completion, and unknown space.
Simulate & learn
Move from the current scene-specific integration to reusable simulation-ready exports with checked scale, collision geometry, and physical parameters. These can support task rehearsal, scenario variation, and training; reduced policy sim-to-real gap remains a separate question to test.
A full mesh is not required for every retrieval task: Memory Over Maps uses posed RGB-D keyframes for on-demand localization, while 3D-Mem and task-oriented scene graphs explore complementary memory and representation choices. Our proposed direction is hybrid: preserve visual evidence, refine task-relevant geometry, and build editable simulation assets where they add value. The next question is when an agent has enough evidence to act—and when it should observe again and update its map.
The immediate agenda is higher reconstruction efficiency and accuracy, reusable simulation-ready scene generation, and deeper integration with phygital agents. Together, these steps connect the reconstruction agent’s modeling and revision loop to the downstream agent’s navigation, memory, and interaction needs.
One synthetic scene and one engineering run per route do not establish general superiority or statistical significance.
The four routes are complete systems, not a strict single-variable ablation; M2 versus M3 cannot be attributed to IMU alone.
M4 uses GT camera poses and is a diagnostic route, not a deployable baseline or a theoretical upper bound.
M1 model-space metrics rely on GT-assisted Sim(3) and must not be ranked as native metric recovery.
The 100-view appearance and depth sets contain modeling RGB inputs; they are not a held-out novel-view benchmark.
Appearance metrics mix geometry, material, illumination, and rendering error. M4 revision 4 has not received a complete independent visual re-review.
The reception-desk demo is map-based simulation execution, not a controlled comparison of navigation across M1–M4. Routes are predefined and localization uses simulator pose; see the demo protocol for scope.
Multimodal memory, long-term map maintenance, and reduced policy sim-to-real gap are research goals, not measured outcomes of the four-robot demo. The Office Café corner observation has no matched-view quantitative GT measurement here.
Historical downstream documentation describes the same M4 Blend hash using a different depth frontend. The frozen tables and assets are retained unchanged; their upstream naming needs reconciliation before attributing historical task outcomes to DA3. See the provenance note.
Every table is generated from frozen assets. The public manifests expose model hashes, registrations, fixed views, and the exact seeded evaluation indices.