Modern 3D tools need precise parameters, yet users often give vague instructions. We introduce CLARE, a clarification-aware, self-evolving 3D agent that resolves this intent asymmetry through strategic dialogue before tool execution. With four specialized roles, CLARE covers five domains—text-to-3D, single-/multi-view reconstruction, point cloud editing, and post-processing—and improves its clarification policy via simulated multi-turn interactions and a Multi-turn Reward. On 3D-Clarify (620 scenarios with ambiguity, missing info, and mistakes), CLARE reaches 60.40% / 43.34% success on single-/multi-step tasks, more than doubling prior baselines.
Modern 3D toolchains require precise, executable parameters, while users often express their goals through vague or underspecified instructions. CLARE addresses this intent asymmetry with a clarification gate that intercepts ambiguity and resolves missing or mistaken details through strategic dialogue before expensive 3D tools are invoked.
CLARE separates its pipeline into four specialized cognitive roles and supports five 3D task domains: text-to-3D generation, single-view reconstruction, multi-view reconstruction, point cloud editing, and post-processing. Its clarification policy self-evolves through simulated multi-turn interactions optimized with a Multi-turn Reward. Evaluation on 3D-Clarify, a benchmark of 620 interaction scenarios, yields success rates of 60.40% on single-step tasks and 43.34% on multi-step tasks.