Soft-Thinking. Official implementation of the NeurIPS 2025 paper "Soft Thinking: Unlocking the Reasoning Potential of LLMs in Continuous Concept Space"
347HarnessAudit. Official codebase for the paper "Auditing Agent Harness Safety"
51arts. Python
23WorldMemArena. Official codebase for the paper "WorldMemArena: Evaluating Multimodal Agent Memory Through Action–World Interaction"
23Length-Value-Model. Official implementation of the paper "Length Value Model: Pretraining Value Model for Scalable Length Prediction and Control"
11SAFEGROUND. SafeGround: Know When to Trust GUI Grounding Models via Uncertainty Calibration
9