Tutorials Self-Host Kimi K3 Open Weights With vLLM
Run Moonshot's 2.8T open-weight model on your own GPUs with vLLM and MXFP4.
How-to content for builders, indie hackers, and AI engineers. Less theory, more shipped code.
Tutorials Run Moonshot's 2.8T open-weight model on your own GPUs with vLLM and MXFP4.
Tutorials Meituan's 1.6T open MoE topped OpenRouter as 'Owl Alpha.' Call it in Python with the OpenAI SDK.
Tutorials Build a frugal tool-calling coding agent on NVIDIA's open Nemotron 3 Nano via OpenRouter in Python.
Tutorials Wire NVIDIA's open 550B MoE into a Python tool-calling loop for long-running agents.
Tutorials Build a cost-aware GLM-5.2 agent that routes thinking effort per task and calls tools.
Tutorials Drive Moonshot's open-weight coding model through a real tool-calling loop in Python.
Machine Learning Run Google's open diffusion LLM with Transformers and learn why it decodes text in parallel.
Tutorials Run Google's open Gemma 4 locally with Ollama and wire up real function calling for an agent.
Tutorials Build a multi-step tool-calling agent on Moonshot's open-weight Kimi K2.6 model.
Tutorials Use Zhipu's GLM-4.7 through the OpenAI SDK to build a tool-calling coding assistant for pennies.
Mobile Open your app from any https URL with Expo Router. iOS + Android setup that actually works.
Tutorials Block agent attacks in <0.1ms with Microsoft's open-source runtime governance toolkit.