WorldSolver: Can LLM Agents Simulate the Physical Dynamics via Solver Generation?
Siru Jiang, Yongzhe Lyu, Shuo Lu, Yubin Wang, Yuxiang Zhang, Yue Liao, Bin Wang, Jian Liang, Tieniu Tan
- Published
- Oct 6, 2026 — 17:26 UTC
Problem
The paper addresses a notable gap in the capabilities of large language model (LLM) agents, specifically their ability to generate solvers for physics simulations. This gap is critical as existing LLMs struggle to produce executable solvers that accurately simulate physical dynamics. The work is presented as a preprint and has not undergone peer review, indicating that the findings are preliminary and subject to further validation.
Method
The authors propose a framework named WorldSolver, which evaluates LLMs on their ability to generate solvers for 168 simulation tasks derived from 61 computer graphics papers across seven distinct physical domains. The evaluation of the generated solvers is conducted along three dimensions:
- Execution Checks: This dimension assesses whether the generated solvers can be executed successfully without errors.
- Visual Fidelity: This measures how well the generated solvers reproduce the intended dynamic behavior visually, ensuring that the output aligns with expected visual outcomes.
- Physical Plausibility: This dimension verifies that the dynamics produced by the solvers adhere to established physics principles, ensuring that the simulations are not only visually accurate but also physically sound.
Results
The evaluation results indicate that the WorldSolver framework achieves a GPT-5.6-Sol Score of 48.7% and a Claude-Opus-5 Score of 46.7% when compared to other evaluated agents. These scores reflect the performance of the LLMs in generating solvers that meet the criteria outlined in the evaluation dimensions, although the paper does not provide further details on the specific baselines or metrics used for comparison.
Limitations
The authors acknowledge several limitations in their approach. A primary challenge is the difficulty in producing executable solvers that consistently perform as intended. Additionally, there are significant hurdles in achieving both visual and physical correctness in the simulations generated by the LLMs. These limitations highlight the ongoing challenges in leveraging LLMs for complex tasks such as physics simulation.
Why it matters
The implications of this work are substantial for downstream applications in computer graphics and simulation. By advancing the capability of LLMs to generate solvers for physics simulations, this research opens avenues for more automated and efficient simulation processes. It also sets the stage for future research to enhance the accuracy and reliability of LLM-generated solvers, potentially transforming how simulations are developed and executed in various domains.
By Turing Wire Research Desk · Oct 6, 2026 · How we work →
Summarised from the paper by the Turing Wire Research Desk. The full paper has the complete methods and results.
Source: arXiv cs.AI
