Pricing: free local AI, hosted GPUs from €19 a month
Free on your hardware. Effortless on ours. Self-host everything for €0, or let LU Labs run the GPUs so you don’t have to.
Flash chat: published daily allowance
Hosted includes 11 Flash models without credits; Pro and Max include 12. Unlimited prompt counts, up to 500,000 combined input and output tokens per account per UTC day in app sessions. All 47 chat models remain available on every plan. GLM 5.3 Flash uses credits on Hosted and is included in the daily allowance on Pro and Max. API keys and accounts without an active subscription always use credits.
App sessions share 500,000 input and output tokens per account per UTC day across the Flash models included on your plan without spending credits. The allowance resets at 00:00 UTC. It is a benefit of an active paid plan, not a free tier: an account without an active plan uses credits for these models, whatever it bought before. API key requests always use credits.
One free request can run at a time per account. A second concurrent request is refused, not silently billed. A request whose token budget exceeds the remaining allowance uses normal credit pricing, with a notice in the app.
We reserve the input estimate and output limit before generation. Without an explicit output limit, free requests allow up to 8,192 output tokens. Each free request lasts at most 240 seconds. Final provider usage settles the reservation. Failed or interrupted requests without final usage keep their reservation. This is a daily allowance, not unlimited usage.
Current model lists and exact monthly image and video budgets →
One app, your choice of where to run
- Every model on every plan.
- Chat, code, image, video and voice in one subscription, one app.
- Runs the same open models on your own machine and in the cloud.
- No five hour windows.
Hosted model access is shared by paid plans and credit packs. Local execution requires compatible hardware and downloaded models; it does not include hosted compute for free. Monthly credit grants, daily Flash allowance and the request limits below still apply. This is not an unlimited-use promise.
Refunds and withdrawal
Our current terms provide EU consumers with withdrawal within 14 days of purchase without giving reasons and a full refund. Contact hello@lu-labs.ai. Subscription cancellation and a refund request are separate actions.
How many images or clips does a budget buy?
Choose a credit pack or a plan, then an image or video model. Every result uses the published catalog and credit rules.
450,000 credits in one shared wallet.
Estimates spend the entire budget on this selection, not on all activities together. Chat draws on the same wallet at each model's published credit rate and is not part of this estimate. Media uses the standard image or 5-second clip rate, without extra operations. Longer clips and other operations can cost more. Plan video limits are included. Annual plans still grant credits monthly. Flash daily usage is separate and is not added to this estimate. Catalog prices can change. This is an estimate, not a guaranteed number of completed generations.
Request limits, not hidden windows
- Chat: 60 requests per minute per account.
- Media job submissions: 30 requests per minute per account.
- Media uploads: 60 requests per minute per account.
- Subscription checkout: 5 attempts per hour and 12 per day per account.
- Credit-pack checkout: a separate 5 attempts per hour and 12 per day per account.
These are request counters, not guaranteed completed generations. They use windows starting with the first request and count requests that reach the limiter, including later validation failures. A rate-limit response includes a Retry-After header.
These burst counters currently live in each server process. Restarts reset them, and multiple server instances do not share them. They are not a global concurrency or availability guarantee. Provider capacity and wallet checks can also refuse a request. The separate Flash daily allowance and one-request limit above are enforced in the database across instances.
Chat model capabilities
Every model on every plan. All 47 chat models are available on paid plans and credit packs. Context and quantization below are from the public DeepInfra model list, checked 2026-09-11. Missing quantization is not a claim of full precision.
Tool transport, image input and reasoning controls describe the LU catalog, not a new live capability test. Provider-listed context is not a guarantee of a full-length prompt: input and output share context, and clients may trim history. No full-native-context promise is made here. Follow each model link for the provider's current information.
| Model and source | Provider-listed context | Provider quantization | Tool calling in LU | Image input in LU | Reasoning controls in LU |
|---|---|---|---|---|---|
| Llama 3.1 8B Turbo | 131,072 | fp8 | Native tools parameter | No | No reasoning control |
| Ling 3.0 flash | 131,072 | Not stated by provider | Native tools parameter | No | low, medium, high; default high |
| Qwen3 30B A3B | 40,960 | fp8 | Native tools parameter | No | low, medium, high; default high |
| Gemma 4 26B | 262,144 | fp8 | Native tools parameter | Yes | low, medium, high; default high |
| Qwen 3.6 35B A3B | 262,144 | fp8 | Native tools parameter | Yes | low, medium, high; default high |
| GLM 5.3 Flash | 1,048,576 | fp4 | Native tools parameter | Yes | low, medium, high, max; default high; no off switch |
| Lunaris 8B | 8,192 | fp8 | LU prompt translation | No | No reasoning control |
| MythoMax 13B | 4,096 | fp16 | LU prompt translation | No | No reasoning control |
| Hermes 3 70B | 131,072 | fp8 | LU prompt translation | No | No reasoning control |
| Euryale 70B | 131,072 | fp8 | LU prompt translation | No | No reasoning control |
| gpt-oss 120B | 131,072 | bfloat16 | Native tools parameter | No | low, medium, high; default high |
| DeepSeek V3.2 | 163,840 | fp4 | Native tools parameter | No | low, medium, high; default high |
| Hermes 3 405B | 131,072 | fp8 | LU prompt translation | No | No reasoning control |
| Qwen3 Coder 480B | 262,144 | fp4 | Native tools parameter | No | No reasoning control |
| Kimi K3 | 1,048,576 | Not stated by provider | Native tools parameter | Yes | low, medium, high; default high |
| DeepSeek V3.1 | 163,840 | fp4 | Native tools parameter | No | low, medium, high; default high |
| DeepSeek V4 Flash 0731 | 1,048,576 | fp8 | Native tools parameter | No | low, medium, high; default high |
| DeepSeek V4.1 Flash | 1,048,576 | fp8 | Native tools parameter | Yes | low, medium, high; default high; no off switch |
| DeepSeek V4 Pro 0813 | 1,048,576 | fp8 | Native tools parameter | No | low, medium, high; default high |
| DeepSeek R1 | 163,840 | fp4 | Native tools parameter | No | low, medium, high; default high |
| Qwen3 32B | 40,960 | fp8 | Native tools parameter | No | low, medium, high; default high |
| Qwen3 235B A22B | 262,144 | fp8 | Native tools parameter | No | No reasoning control |
| Qwen 3.5 9B | 262,144 | bfloat16 | Native tools parameter | Yes | low, medium, high; default high |
| Qwen 3.5 35B A3B | 262,144 | fp8 | Native tools parameter | Yes | low, medium, high; default high |
| Qwen 3.5 397B A17B | 262,144 | fp8 | Native tools parameter | Yes | low, medium, high; default high |
| Qwen 3.6 27B | 262,144 | fp8 | Native tools parameter | Yes | low, medium, high; default high |
| Qwen3 VL 30B | 262,144 | fp8 | Native tools parameter | Yes | No reasoning control |
| Qwen3 VL 235B | 262,144 | fp8 | Native tools parameter | Yes | No reasoning control |
| Qwen 3.8 27B | 262,144 | Not stated by provider | Native tools parameter | Yes | low, medium; default medium |
| Qwen 3.8 Max | 256,000 | Not stated by provider | Native tools parameter | No | low, medium, high; default high |
| Qwen 3.8 A95B | 262,144 | fp4 | Native tools parameter | No | low, medium, high; default high |
| Llama 3.3 70B Turbo | 131,072 | fp8 | Native tools parameter | No | No reasoning control |
| Llama 4 Maverick | 1,048,576 | fp8 | LU prompt translation | Yes | No reasoning control |
| Llama 4 Scout | 327,680 | fp8 | Native tools parameter | Yes | No reasoning control |
| Gemma 4 31B Turbo | 262,144 | fp4 | Native tools parameter | Yes | low, medium, high; default high |
| GLM 4.7 | 202,752 | fp4 | Native tools parameter | No | low, medium, high; default high |
| GLM 5 | 202,752 | fp4 | Native tools parameter | No | low, medium, high; default high |
| GLM 5.1 | 202,752 | fp4 | Native tools parameter | No | low, medium, high; default high |
| GLM 5.2 | 1,048,576 | fp4 | Native tools parameter | No | low, medium, high; default high |
| GLM 5.3 | 1,048,576 | fp4 | Native tools parameter | No | low, medium, high, max; default high; no off switch |
| Nemotron 3 Super 120B | 262,144 | bfloat16 | Native tools parameter | No | low, medium, high; default high |
| gpt-oss 20B | 131,072 | bfloat16 | Native tools parameter | No | low, medium, high; default high |
| Kimi K2.6 | 262,144 | fp4 | Native tools parameter | Yes | low, medium, high; default high |
| Kimi K2.7 Code | 262,144 | fp4 | Native tools parameter | Yes | low, medium, high; default high |
| MiniMax M2.7 | 196,608 | fp8 | Native tools parameter | No | low, medium, high; default high |
| MiniMax M3 | 524,288 | fp8 | Native tools parameter | Yes | low, medium, high; default high |
| Mistral Small 3.2 24B | 128,000 | fp8 | Native tools parameter | Yes | No reasoning control |