Long Horizon Terminal Benchmark with Dense Reward Grading
-
Updated
Aug 27, 2026 - Python
Long Horizon Terminal Benchmark with Dense Reward Grading
Zenith: a continuous-improvement harness for long-running agent tasks. Turns Claude Code, Codex, or Hermes into a multi-agent mission orchestrator via MCP/ACP.
EvoX Genesis is an autonomous system for long-horizon software evolution that recursively builds, continues, and transforms complex software from high-level objectives.
VLM-RL Hierarchical Loco-Manupilation For Long-Horizon Tasks With G1 robot in Isaac Lab/Sim
Official code for Behavior-Skill, a fine-grained skill dataset and evaluation benchmark for Vision-Language-Action policies in long-horizon mobile manipulation tasks.
Eureka — task-conditioned Meta-Agent orchestration for scientific discovery, developed by ManXis.
This package supports global planning and meta-control for AI agents tackling complex, long-horizon tasks, reducing costly strategic mistakes and blind trial-and-error search.
A reliability layer for long-running AI agents - progress survives crashes, context loss, and executor changes without relying on chat history. Ships as a normative spec, an installable agent skill, and a stdlib-only Python reference runtime.
A scenario library and testing framework for GenAI alignment — objective drift, robustness, and agentic/enterprise risk — tested against Azure OpenAI and Claude, with audit-ready findings mapped to governance frameworks.
Durable long-horizon task runtime for DeepSeek Harness — plan, confirm, track, and resume AI work across sessions and interruptions.
Long-horizon AI harness you actually own: one static Go binary, one persona, any OpenAI-compat LLM (Ollama, Gemini, Claude, ChatGPT), MCP tools, chat via Telegram/Discord/Slack. Outbound-only — no dashboard, no config UI, no open ports, ever. Hardened so small local models finish tool calls.
A benchmark for evaluating coding-agent memory systems through long-horizon personal worklog tasks
Run multi-day autonomous engineering campaigns with coding-agent goals (Codex /goal) — goal-spec + state-file governance templates, adversarial convergence gates, and a demo API + k6 harness to try the pattern locally
Self-improving long-horizon LLM agent — ChromaDB strategy memory + failure analysis, Grok-4 teacher labels → QLoRA-distilled LLaMA-3.2-1B student. 90% on Tau Bench, 95% inference cost reduction.
Long-horizon task management plugin for Claude Code: 辅助增强 Claude Code 的长程任务能力
To associate your repository with the long-horizon-tasks topic, visit your repo's landing page and select "manage topics."