🏃 Sovereign Agent Evaluation Framework - Zero cloud dependencies. Local-only, cryptographically signed benchmarks for AI agents. npm run demo = instant evals.
-
Updated
Mar 23, 2026 - TypeScript
🏃 Sovereign Agent Evaluation Framework - Zero cloud dependencies. Local-only, cryptographically signed benchmarks for AI agents. npm run demo = instant evals.
TraceOS standardizes AI experiments into reproducible, searchable, and comparable assets. One command runs experiments, generates reports, and produces structured analysis: capability vectors, failure taxonomy, and recommendations. Every run is tracked, traceable, and comparable. Built on ABC-130K (amazon-far/abc). Apache 2.0.
Adaptive multi-star candidate ranking system using deterministic scoring + relevance feedback with bias-aware filtering.
An end-to-end Machine Learning project featuring a modular pipeline, configuration-driven workflows, MLflow experiment tracking, DagsHub integration, and a Flask web interface, following industry-standard MLOps practices.
σFlow-PDE: A drop-in H-Bar training engine that escapes the σ-trap in neural PDE solvers via live σ/δ/α ODE integration, autonomous phase curriculum, and auto-falsification.
Production-ready ML system for credit default prediction on transactional data (458k clients). Features end-to-end pipeline: 1,158 engineered features, ablation & Top-500 pruning, LightGBM HPO (AMEX 0.791, Gini 0.923), multi-seed stability, CLI batch inference (20.1s/458k), and 267/267 automated tests.
Enterprise digital laboratory for machine learning that organizes the full ML development lifecycle in a managed, reproducible form within a single operational context with common execution, security, and audit rules.
A reproducible visual-attribute verification framework combining group-disjoint evaluation, audited LoRA controls, calibration analysis, and CI-backed evidence contracts.
Deterministic job decision engine that scores opportunities using a transparent, testable formula and logs every decision with full traceability. Hybrid 5-signal scoring with a bounded LLM reasoning layer. Same input gives the same output on the LLM-free path, verified to 1e-9 in local and CI runs.
Add a description, image, and links to the reproducible-ml topic page so that developers can more easily learn about it.
To associate your repository with the reproducible-ml topic, visit your repo's landing page and select "manage topics."