You signed in with another tab or window. Reload to refresh your session.You signed out in another tab or window. Reload to refresh your session.You switched accounts on another tab or window. Reload to refresh your session.Dismiss alert
🖨️ Automated scanner document processor with AI-powered naming and WebDav integration. Receives scans via FTP, extracts text using Vision AI, generates intelligent filenames with Ollama AI, and uploads to your cloud storage.
AI-powered browser automation agent using a dual-LLM architecture. The orchestrator (qwen3-vl-32k) creates execution plans from screenshots, while the executor (llama3.1:8b) translates steps into browser actions using an accessibility tree for reliable element selection. Local, private, powered by Ollama.
AI video editing skill that watches your footage before cutting. Vision model analyzes every frame, scores usability, drops bad takes, reorders by narrative, then renders with ffmpeg. For Claude Code, Cola, OpenClaw. 会看画面的 AI 自动剪辑 Skill——先用视觉模型看懂素材再决定剪哪几秒。
Локальная мульти-агентная система на self-hosted LLM (Ollama). 3 канала (Telegram/консоль/MAX), 4 типа моделей (LLM/embed/vision/voice), мультимодальность (текст/PDF/фото/голос), роли Planner/Executor/Critic с рефлексией. Инструменты: почта, Яндекс.Диск, web, OCR, cron-планировщик. Семантическая память (sqlite-vec). Coverage 88%. Без облачных API.
This is a fork of SpaceInvaderOne's repo to fix some issues I had with the software until he pulls the changes or fixes them himself. It also allows for integration to my gallery app Eyeris. github.com/vonhex/eyeris
Codex plugin that lets Codex read images via a pure vision model (base64 -> vision API -> back to main model). 让 Codex 通过纯视觉模型查看图片的插件:图片转 base64 后调用视觉模型,识别结果返回给主模型。
Next-gen AI Optical Music Recognition (OMR) platform. Convert sheet music images into playable ABC notation instantly using Google Gemini 3 Pro Vision. Built with React 19, TypeScript, and Tailwind.