Skip to content

Latest commit

 

History

History
65 lines (51 loc) · 2.77 KB

File metadata and controls

65 lines (51 loc) · 2.77 KB

Sprint 10 Report — Local LLM Research & Optimization Strategy

Datum: 2026-05-13 Status: ✅ COMPLETE

1. Erledigte Tracks (8/8)

Track Thema Datei Größe
A Beste lokale Modelle für Story Writing track_a_modelle_recherche.md 17.5 KB
B Training / Finetuning Möglichkeiten track_b_training_methods_comparison.md 17.9 KB
C Dataset-Strategien track_c_dataset_strategies.md 14.8 KB
D RAG vs Finetune track_d_rag_vs_finetune.md (in TRACK_D_RAG_VS_FINETUNE_ANALYSIS.md) 23.6 KB
E Multi-Model Orchestration research/track_e_multi_model_orchestration.md 21.4 KB
F Lokale Infrastruktur track_f_infrastructure_comparison.md 24.0 KB
G Memory / Long Context TRACK_G_MEMORY_LONG_CONTEXT_RESEARCH.md 20.1 KB
H Praktische Gesamtempfehlung docs/local_llm_strategy.md 33.5 KB

2. Kernerkenntnisse

Top-Modelle für RTX 3060 12GB

Rang Modell VRAM Story-Qualität
🥇 Gemma 3 27B (IQ3_M) 11 GB 4.3/5 (übertrifft 70B-Modelle)
🥈 Qwen 3.6 27B (IQ3_M) 10.5 GB Modernste Hybrid-Attention
🥉 Phi-4 14B (Q4_K_M) 9 GB Ideal für Editing/Summaries

Bestes Finetuning: QLoRA via Unsloth

  • 99% der Full-FT Performance auf 12 GB VRAM
  • 300–500 handkuratierte Roman-Auszüge > 50.000 generische Samples
  • Raw Continuation Format (kein Chat-Template)
  • Kosten: €0 (nur Strom)

Dataset-Strategie

  • Gutenberg + Fanfiction als Basis
  • 5.000 Beispiele = Sweet Spot für QLoRA
  • Anti-Slop-Filter (Regex + Perplexity + Diversity)
  • Deutsche Daten: größte Lücke

Infrastruktur

  • Primär: Ollama + Open WebUI (RAG, Multi-Model, Tool Calling)
  • Alternative: llama.cpp-server (+20% Tokens/s als Drop-In)

3. Drei empfohlene Stacks

Stack Kosten/Monat Qualität Für wen
LOW-END $0 6.5/10 Hobby-Autoren, Local-First
MID-TIER $1–5 8.2/10 Semi-professionell, hybrid
HIGH-END $20–50 9.2/10 Professionell, Cloud-primary

4. TOP-5 Verbesserungsvorschläge

# Vorschlag Impact Aufwand
1 SummaryBufferMemory + Reflection 🔥 Extrem 5–7 Tage
2 QLoRA-Finetuning für Stil 🔥 Sehr hoch 2–5 Tage
3 Pipeline-Orchestrierung 🔥 Sehr hoch 3–5 Tage
4 RAG-Optimierungen 🔥 Hoch 1–2 Tage
5 Memory Consolidation Service 🔥 Hoch 5–10 Tage

5. Gesamtdokument

F:\bookgenerator_hermes\docs\local_llm_strategy.md (33.5 KB) — Enthält alle 3 Stacks, Entscheidungsbaum, Installationsanleitung, Pipeline-Diagramme, Sprint-10–14 Roadmap