Open-weight shortlist
Four language alternatives and nine media models worth self-hosting.
This is a filtered list, not a copy of a hosted catalogue. Every media entry has released weights, a credible local or workstation route, and creator-published quality or benchmark evidence. Hosted-only products, LoRAs, duplicate endpoints and utility wrappers stay out.
16 GB · 14B Q4 9.1 GBMicrosoft Phi-4
A compact MIT-licensed reasoning model. Microsoft reports MMLU 84.8, GPQA 56.1 and HumanEval 82.6 for the 14B base instruction model; reasoning, mini and multimodal Phi-4 variants also ship.
Sizes14Bminimultimodalreasoning
8-32 GB · Apache 2.0IBM Granite 4
Enterprise-oriented models with governance disclosures, hybrid Mamba/Transformer variants and compact footprints. Shipped language sizes include Micro, H-Micro, H-Tiny, H-Small and dense 8B; evaluate the exact card because scores vary by variant.
SizesMicroH-MicroH-TinyH-Small8B
16 / 32 GB · fully open researchAI2 OLMo 3
A rare option with training code, data, checkpoints and detailed recipes. The family ships 7B and 32B Base, Instruct and Think variants with 65,536-token context; use 7B Q4/Q8 at 16 GB and 32B Q4 at 24-32 GB.
Sizes7B32BBaseInstructThink
Estonian specialist · prototypeTartuNLP EstLLM 8B
A locally testable Estonian-focused research checkpoint and evaluation suite. The authors clearly label it an early prototype with 4K context and no multi-turn chat; use it as a benchmark and fine-tuning reference, not an automatic production default.
Size8BBF16 checkpointcommunity quants
13 / 29 GB VRAM · local image specialistFLUX.2 Klein 4B + 9B
The 4B and 9B checkpoints combine text-to-image generation and multi-reference editing in one four-step model. Black Forest Labs says the Apache-2.0 4B model fits about 13 GB VRAM on an RTX 3090 or 4070-class GPU; the stronger 9B model fits about 29 GB and uses the FLUX Non-Commercial License. Both are actual WaveSpeed endpoints, but their released weights make local use credible without the API.
Creator-published fit4B · about 13 GB VRAM9B · about 29 GB VRAM4 inference stepsText + multi-reference editing
Under 16 GB VRAM · Apache 2.0Z-Image Turbo
Tongyi-MAI's 6B image model targets photorealism and bilingual text rendering in a much smaller package than the server leaders. The creator says it runs on consumer GPUs with under 16 GB VRAM and publishes the code and weights. That makes the Turbo checkpoint a genuine local alternative, not merely a WaveSpeed product label.
Creator-published fit6B parametersUnder 16 GB VRAMText-to-imageEnglish + Chinese text
14 GB VRAM with offload · videoHunyuanVideo 1.5
Tencent's 8.3B text-to-video and image-to-video release is the strongest current lightweight video candidate in this catalogue. Tencent reports state-of-the-art open-model quality and motion coherence, publishes training and inference code, and documents a 14 GB minimum with model offloading. The step-distilled 480p image-to-video checkpoint runs in 8 or 12 steps and was measured by Tencent at under 75 seconds on an RTX 4090.
Creator-published fit8.3B parameters14 GB VRAM minimumText + image to video480p distilled + 720p
24 GB VRAM · 720p videoWan 2.2 TI2V 5B
The 5B Wan checkpoint handles both text-to-video and image-to-video at 720p and 24 fps on one 24 GB RTX 4090 with offloading. Alibaba describes it as one of the fastest open 720p models and publishes the weights and inference code. The larger A14B route belongs on the server page; this 5B checkpoint is the consumer-workstation cut.
Creator-published fit5B parameters24 GB VRAM720p24 fps
8-24 GB VRAM with offload · music generationMiniMax Music 3
MiniMax Music 3 generates complete songs of up to five minutes from lyrics and a detailed music description, including vocals and 32 kHz stereo audio. The official Diffusers path fits the full-precision pipeline in under 24 GB VRAM; automatic CPU offload uses about 22 GB, and layer-by-layer language-model streaming can reduce the GPU requirement to 8 GB at the cost of speed. Inference currently requires CUDA and does not stream the generated audio.
Creator-published fitUnder 24 GB · full precisionAbout 22 GB · CPU offload8 GB · layer streamingUp to 5-minute songs
0.6B / 1.7B · speech specialistQwen3-TTS
Qwen releases 0.6B and 1.7B speech models for custom voices, voice design and cloning. On Qwen's Seed-TTS evaluation, the 12 Hz 1.7B Base checkpoint reports 0.77 Chinese and 1.24 English word error rates, ahead of the commercial and open baselines listed in that table. The official project does not state a minimum VRAM, so treat the small parameter count as a starting signal and measure the complete audio stack on the target GPU.
Released sizes0.6B1.7BCustom voiceVoice design + cloning
4-24 GB VRAM · music generationACE-Step 1.5
ACE-Step is the clearest local music-generation addition in the hosted catalogue. The standard model can run with under 4 GB VRAM, while its 4B XL decoder needs at least 12 GB with offload and recommends 20 GB. The creators report under ten seconds per full song on an RTX 3090 and publish benchmark, profiling and GPU-tier tools, including Mac, AMD, Intel and CUDA paths.
Creator-published fitStandard: under 4 GB VRAMXL: 12 GB with offloadXL: 20 GB recommendedUp to 10-minute songs
10-29 GB VRAM · 3D assetsHunyuan3D 2.1
Tencent releases the full 3.3B shape and 2B PBR texture pipeline, including training code. Its creator evaluation leads the listed open and closed baselines on shape following and textured-asset metrics. Hardware scales by task: 10 GB for shape, 21 GB for texture, or 29 GB for the complete pipeline. Note that its community licence excludes the EU, UK and South Korea.
Creator-published fitShape: 3.3B + 10 GB VRAMTexture: 2B + 21 GB VRAMComplete: 29 GB VRAMPBR materials
3.44 GB checkpoint · segmentationSAM 3.1
Meta's current Segment Anything release handles text-prompted image segmentation, video segmentation and multi-object tracking. The official checkpoint is about 3.44 GB. On Meta's public tables, SAM 3.1 improves six of seven video-object-segmentation benchmarks over SAM 3 and adds an approximately sevenfold 128-object speedup on one H100 test. The official project does not publish a consumer-GPU minimum, so measure your prompt count, image size and video workload before promising a device tier.
Released route3.44 GB checkpointImage masksVideo trackingText + visual prompts