Productive, portable, and performant GPU programming in Python.
-
Updated
Jul 6, 2026 - C++
Productive, portable, and performant GPU programming in Python.
High-performance large-scale embedding acceleration for JAX on Google TPU SparseCores.
Leveraging Taichi Lang to customize brain dynamics operators.
Improving SpGEMM Performance Through Matrix Reordering and Cluster-wise Computation [SC'25]
Sparse Lie algebra engine for G₂, F₄, E₆, E₇, E₈ — 913× compression, lattice gauge theory, equivariant GNN layers. pip install dhl-mm
Relation-based language modeling, RelationLex tokenization, stateful decode, and fused Triton kernels
CPU-native inference runtime. Local-propagation paradigm: the active region pays the cost, not the field. Bit-exact across architectures. Validated for streaming anomaly detection and audio VAD.
Research code for testing whether causal token surprisal can guide adaptive computation, sparse refinement, and learned compute allocation in byte-level language models.
Add a description, image, and links to the sparse-computation topic page so that developers can more easily learn about it.
To associate your repository with the sparse-computation topic, visit your repo's landing page and select "manage topics."