Custom OpenAI-compatible AI setup
Your own model server, or any other OpenAI-compatible API.
What it does
Anything that speaks the OpenAI chat completions API works here: Ollama or LM Studio on your own machine or VPS, vLLM, or a paid provider you already use.
- The very last AI fallback, after every free provider
Get set up
- Run an OpenAI-compatible server, for example Ollama (
ollama serve, thenollama pull llama3.1) - Make it reachable from the internet over HTTPS, protected by a key
- In Settings → AI, set Custom AI base URL (ending in
/v1), Custom AI model and, if needed, Custom AI key - Click Check next to it on the Jobs & providers page
Where it goes
In the dashboard, open Settings and find the AI section. Secret values are encrypted before they're saved and only their last four characters are ever shown again.
LLM_BASE_URLLLM_MODELLLM_API_KEYStaying free
It is tried last, so a paid endpoint only gets the calls every free provider turned down. It still has to pass the AI quality check before it writes anything.
How the quota guard works →Larva tip
Small local models (under ~8B parameters) often fail the quality check. That's the check doing its job: their drafts read like drafts.
Fallback order
For AI, Larvabot tries these in order and skips any that aren't set up, are out of free quota or are busy. AI providers also have to pass the quality check on the Jobs & providers page.
- 1Google Gemini
- 2Groq
- 3OpenRouter (:free models)
- 4Cloudflare Workers AI
- 5Cerebras
- 6Mistral
- 7Custom OpenAI-compatible (e.g. Ollama)