Krasis is a Hybrid LLM runtime which focuses on efficient running of larger models on consumer grade VRAM limited hardware
-
Updated
Aug 15, 2026 - C++
Krasis is a Hybrid LLM runtime which focuses on efficient running of larger models on consumer grade VRAM limited hardware
Interactive 3D visualization platform for exploring transformer architectures, tensors, and real-time LLM inference.
GGUF Quantization support for native ComfyUI models
Local inference platform for K/IQ-quant GGUF models on Apple Silicon
Android inference engine running 20B+ parameter LLMs on 4GB-8GB RAM devices. Features proprietary Layer-by-Layer (LBL) streaming, zero-copy mmap loading, and native C++/Kotlin architecture.
A tiered-memory system design for workloads that don't fit in RAM: measure the working set, pin the hot tier, stream the cold tier from flash. Ships the residency calculator, measurement harnesses, and the build recipes behind it. Predictions validated against public benchmarks.
Running GGUF-Format ML Models in Pure Go
Splinter cohabitates inference and semantic governance in L3 cache and memory lanes, while simultaneously providing standardized POSIX-friendly tooling as building blocks on top of the provided library. Splinter is essentially a semantic "breadboard" that can be directly deployed at scale.
Privacy-first Local RAG Server: Chat with PDF & DOCX using GGUF models via llama.cpp and Qdrant. A lightweight, standalone FastAPI server with a clean HTML UI. High-performance, fully offline document intelligence. No Ollama, no cloud, no API keys.
A simple Gradio app for local translation using the GGUF versions of MADLAD-400
Convert and quantize llm models
Emotica AI is a compassionate and therapeutic virtual assistant designed to provide empathetic and supportive conversations. It integrates a local LLaMA model for text generation, a vision model for image captioning, a RAG system for information retrieval, and emotion detection to tailor its responses.
Nectar-X-Studio is a powerful, Local AI-Inferencing application that allows the user download, create, run agents and run large language models on their own machine. With no internet connection required, Nectar ensures privacy-first, high-performance inference using cutting-edge open-source models from Hugging Face, Ollama, and beyond.
Containerized LLM for any use-case big or small
AI tool to help users research using local LLMs and automated web search.
Render is a lightweight, easy-to-use CLI and local server for generating AI images. Run Stable Diffusion (SD 1.5, SDXL) and FLUX models locally with simple commands like pull and run. Powered by Vulkan acceleration and GGUF support for ultra-fast performance. Think Ollama, but for local image generation.
GGUF file format for dotnet
Add a description, image, and links to the gguf-model-support topic page so that developers can more easily learn about it.
To associate your repository with the gguf-model-support topic, visit your repo's landing page and select "manage topics."