Gwenlake API
One OpenAI-compatible API for chat, embeddings, reranking and audio. We deploy open-source models for you on sovereign infrastructure and operate them: the GPUs, the scaling, the upgrades. Your applications only ever see an endpoint.
Four capabilities, one endpoint
A retrieval pipeline needs more than a chat model. Everything it takes is served from the same API, under the same key and the same quota.
Chat & completion
Open-source LLMs served with streaming, tool calling and structured output, through the OpenAI-compatible shape your libraries already speak.
Embeddings
Vectorise documents and queries with the embedding models of your choice, on the same sovereign infrastructure as the rest.
Reranking
Reorder retrieved passages by real relevance, the step that turns a mediocre RAG pipeline into a usable one.
Audio
Transcription for the calls, meetings and recordings that hold as much of your knowledge as your documents do.
No third party in the path
Private inference is not a setting we tick. It is the reason the service exists.
Open-weight models
Deployed and operated on sovereign infrastructure.
Dedicated endpoints
Isolated per organisation and per project.
Nothing trains a model
Your prompts and documents are never used to train anyone’s model.
Zero retention
For the workloads that must leave no payload behind.
Only models we have tested
Nothing runs on this API that we have not evaluated ourselves. Either it comes out of our own research, or it is an open-weight model we have benchmarked, deployed and operated before putting it on the menu.
Our own work
The algorithms and models we develop ourselves, out of the research we keep doing alongside client projects.
Open weights, measured
Open-source and open-weight models are benchmarked on real tasks before they reach the catalogue, and measured again when a new version lands.
No black boxes
Nothing is proxied to a model we cannot inspect, host and explain. If we cannot run it ourselves, it is not on offer.
A catalogue you can follow
The catalogue is configuration, not code: we add, version and retire models without your applications changing a line, and we tell you when one moves.
Point your application at it
If your code already speaks the OpenAI API, moving to a sovereign endpoint is a base URL and a key. Tell us what you are running and we will size it with you.