Local AI — Community
5 articles on Local AI from Ayra ix.
Local AI
LiteLLM on the Homelab: One Proxy, Every Model, Waterfall Routing Included
LiteLLM transforms your homelab into a flexible AI gateway with fallback chains, load balancing, and cost control across local and remote LLMs.
Local AI
Ollama GPU Scheduling: Running Inference and ComfyUI on One RTX Without OOM
Strategies for sharing a single RTX GPU between Ollama LLM inference and ComfyUI Stable Diffusion on the same homelab machine.
Local AI
Small Models at the Edge: When 3B Beats 70B
Not every inference needs a datacenter brain. Small models on device win on latency, privacy, and bills — if you scope them right.
Local AI
Open-Weight Convergence: The Models Are Similar — Now What?
The gap between top open models narrowed in 2026. Your moat was never the weights — here's what actually differentiates.
Local AI
Your GPU Is a Datacenter — Start Treating It Like One
That RTX card in your closet isn't a toy. It's a private inference region with power, cooling, and SLA implications you ignore at your peril.