A course of learning LLM inference serving on Apple Silicon for systems engineers: build a tiny vLLM + Qwen.
skyzh/tiny-llm is engineered as a server, cloud service, or developer package. It is primarily deployed via Docker containers or package managers rather than a single desktop executable.