for developers
Three steps. Your code does not change: any OpenAI client works once it points at the gateway.
get a key
Made in the dashboard after one wallet signature. It is shown once; only a hash is stored, so a lost key is replaced rather than recovered.
open the dashboard keys look likesk-api-…Keys are issued at launch, once the gateway is live.
point your client at the gateway
Two settings. Paste them wherever your tool asks for an OpenAI key.
base urlhttps://holdapi.lol/v1api keysk-api-… the key from step 1yoursOr pick your tool and copy the whole snippet.
Cursor -> Settings -> Models -> OpenAI API Key Base URL https://holdapi.lol/v1 API Key sk-api-... Turn on "Override OpenAI Base URL", then pick a model from the list on this site.
openai compatible · cursorpick a model
Use the id as the model name. Rates are published per million tokens, and the endpoint meters your real usage against them instead of rounding to a plan.
models · read from /v1/models
openai/gpt-6-astraopenaiopenai/gpt-6.1-solopenaiopenai/gpt-6-solopenaiopenai/gpt-6-lunaopenaiopenai/gpt-5.6-solopenaiopenai/gpt-5.6-terraopenaiopenai/gpt-5.6-lunaopenaiopenai/gpt-5.6-sol-proopenaiopenai/gpt-5.6-terra-proopenaiopenai/gpt-5.6-luna-proopenaiopenai/gpt-5.5openaiopenai/gpt-5.5-proopenaiopenai/chat-latestopenaiopenai/gpt-5.4openaiopenai/gpt-5.4-proopenaiopenai/gpt-5.1openaiopenai/gpt-5.2openaiopenai/gpt-5.4-miniopenaiopenai/gpt-5-miniopenaiopenai/gpt-5.4-nanoopenaiopenai/gpt-5.2-proopenaiopenai/gpt-5.3-codexopenaiopenai/gpt-4.1openaiopenai/gpt-4.1-miniopenaiopenai/gpt-4.1-nanoopenaiopenai/gpt-4oopenaiopenai/gpt-4o-miniopenaiopenai/o1openaiopenai/o3openaiopenai/o3-miniopenaiopenai/o4-miniopenaianthropic/claude-haiku-4.5anthropicanthropic/claude-sonnet-5.5anthropicanthropic/claude-sonnet-5anthropicanthropic/claude-sonnet-4.6anthropicanthropic/claude-sonnet-4.5anthropicanthropic/claude-opus-4.5anthropicanthropic/claude-opus-4.7anthropicanthropic/claude-fable-5.1anthropicanthropic/claude-fable-5anthropicanthropic/claude-opus-4.8anthropicanthropic/claude-opus-5anthropicanthropic/claude-opus-5.5anthropicgoogle/gemini-3.1-progooglegoogle/gemini-3-flash-previewgooglegoogle/gemini-3.8-flashgooglegoogle/gemini-3.6-flashgooglegoogle/gemini-3.5-flashgooglegoogle/gemini-2.5-progooglegoogle/gemini-2.5-flashgooglegoogle/gemini-3.5-flash-litegooglegoogle/gemini-3.1-flash-litegooglegoogle/gemini-2.5-flash-litegoogledeepseek/deepseek-v4-flash-vision-expdeepseekdeepseek/deepseek-v4-prodeepseekdeepseek/deepseek-chatdeepseekdeepseek/deepseek-reasonerdeepseekmoonshot/kimi-k3moonshotzai/glm-5.3zaizai/glm-5.3-flashzaizai/glm-5.2zaizai/glm-5.1zaizai/glm-5zaizai/glm-5-turbozaixai/grok-4.3xaixai/grok-build-0.1xaixai/grok-4.7xaixai/grok-4.6xaixai/grok-4.5xaiminimax/minimax-m2.7minimaxminimax/minimax-m3minimaxqwen/qwen3.8-maxqwenqwen/qwen3.7-maxqwenqwen/qwen3.7-plusqwenqwen/qwen3.7-flashqwenqwen/qwen3.8-flashqwentencent/hy4-previewtencentxiaomi/mimo-v2.5xiaomixiaomi/mimo-v2.5-proxiaominvidia/nemotron-3-nano-omni-30b-a3b-reasoningnvidianvidia/nemotron-3.5-lightningnvidianvidia/llama-3.2-11b-visionnvidianvidia/nemotron-3-ultra-550bnvidiacohere/north-mini-codecoherepoolside/laguna-xs-2.1nvidiapoolside/laguna-s-2.1poolsidemistral/mistral-large-4mistralnvidia/muse-glimmer-30bnvidia
A call is charged what the model costs. Each call first holds the most it could cost, then settles to what it used and returns the difference. That is why a long answer can hold more than it spends. If a model is not in this list, the endpoint will not run it.