Long Horizon Terminal Benchmark with Dense Reward Grading
-
Updated
Jul 23, 2026 - Python
Long Horizon Terminal Benchmark with Dense Reward Grading
Zenith: a continuous-improvement harness for long-running agent tasks. Turns Claude Code, Codex, or Hermes into a multi-agent mission orchestrator via MCP/ACP.
VLM-RL Hierarchical Loco-Manupilation For Long-Horizon Tasks With G1 robot in Isaac Lab/Sim
This package supports global planning and meta-control for AI agents tackling complex, long-horizon tasks, reducing costly strategic mistakes and blind trial-and-error search.
Run multi-day autonomous engineering campaigns with coding-agent goals (Codex /goal) — goal-spec + state-file governance templates, adversarial convergence gates, and a demo API + k6 harness to try the pattern locally
Self-improving long-horizon LLM agent — ChromaDB strategy memory + failure analysis, Grok-4 teacher labels → QLoRA-distilled LLaMA-3.2-1B student. 90% on Tau Bench, 95% inference cost reduction.
Add a description, image, and links to the long-horizon-tasks topic page so that developers can more easily learn about it.
To associate your repository with the long-horizon-tasks topic, visit your repo's landing page and select "manage topics."