Historical Windows temporal memory-state research artifact for studying time-bound memory observations, validation limits, and defensive visibility.
-
Updated
May 15, 2026 - Python
Historical Windows temporal memory-state research artifact for studying time-bound memory observations, validation limits, and defensive visibility.
PCBench: Benchmark for Python API parameter compatibility issues
Artifact for "Love, Lies, and Language Models" (USENIX Security '26) — evaluating whether commercial content moderation detects romance-baiting fraud. It largely doesn't.
Code, data, and ontologies for FAOS research papers on ontology-powered enterprise AI agent verification (RA-3 neurosymbolic, RA-6 trust certification).
Reproducibility artifact for the TMLR paper on quantization and cue-conditioned reasoning: frozen trials, 15 generation files, 3,078 automated judgments, human annotations, and analysis code
SymbolicLight V1: spike-gated dual-path language model code, tokenizers, and reproducibility artifacts for 194M and 0.8B releases.
Experimental Python runtime for validation-gated program synthesis and adaptive search: multi-level meta-learning (meta-meta loops), analogical transfer, grammar-mediated expansion, anti-cheat verifiers, sealed evaluations, rollback-sensitive self-modification. Bounded adaptive improvement, not unrestricted recursive self-improvement.
JSON Schema for decision events as governance evidence units in automated decision and real-time risk systems. MIT.
Historical lab-specific prototype for controlled volatile-memory and network capture research.
Thesis artifact: RAMA, a resilient microservices architecture for multi-model AI inference on resource-constrained nodes.
Research artifact for One QK Channel, Many Sources (arXiv:2608.02091)
Evaluation infrastructure for AI systems beyond direct human supervision
Side-channel profiler that detects deceptive intent in LLMs by measuring the computational cost of lying.
Reproducible evaluation of compiled court-form schemas against chunked and induced retrieval on synthetic FW-001 fixtures.
Code, data, and raw model outputs for "Failure Modes in Perturbation-Based Measurement of Language Model Reliability". evaluation/reproduce.py re-derives every reported number offline, no API key required.
Curated code and result summary for world-model inputs in Atari policy experiments.
Reference artifacts for "A Governance-Layer Architecture for Human-Governed AI Workforces" — PostgreSQL schemas, Rule H10, HR-evaluator queries, templates, synthetic ledger, companion v7.0 spec
Research artifact for generating rotated image datasets and evaluating CNN/CyCNN models across rotation-based train-test scenarios.
Synthetic matched-control benchmark for detecting effective override loss in AI-mediated workflows.
Research artifact repository containing the AB Genesis Simulator and benchmark framework fo...
Add a description, image, and links to the research-artifact topic page so that developers can more easily learn about it.
To associate your repository with the research-artifact topic, visit your repo's landing page and select "manage topics."