Gwenlake API

One OpenAI-compatible API for chat, embeddings, reranking and audio. We deploy open-source models for you on sovereign infrastructure and operate them: the GPUs, the scaling, the upgrades. Your applications only ever see an endpoint.

Four capabilities, one endpoint

A retrieval pipeline needs more than a chat model. Everything it takes is served from the same API, under the same key and the same quota.

  • Chat & completion

    Open-source LLMs served with streaming, tool calling and structured output, through the OpenAI-compatible shape your libraries already speak.

  • Embeddings

    Vectorise documents and queries with the embedding models of your choice, on the same sovereign infrastructure as the rest.

  • Reranking

    Reorder retrieved passages by real relevance, the step that turns a mediocre RAG pipeline into a usable one.

  • Audio

    Transcription for the calls, meetings and recordings that hold as much of your knowledge as your documents do.

No third party in the path

Private inference is not a setting we tick. It is the reason the service exists.

  • Open-weight models

    Deployed and operated on sovereign infrastructure.

  • Dedicated endpoints

    Isolated per organisation and per project.

  • Nothing trains a model

    Your prompts and documents are never used to train anyone’s model.

  • Zero retention

    For the workloads that must leave no payload behind.

Only models we have tested

Nothing runs on this API that we have not evaluated ourselves. Either it comes out of our own research, or it is an open-weight model we have benchmarked, deployed and operated before putting it on the menu.

  • Our own work

    The algorithms and models we develop ourselves, out of the research we keep doing alongside client projects.

  • Open weights, measured

    Open-source and open-weight models are benchmarked on real tasks before they reach the catalogue, and measured again when a new version lands.

  • No black boxes

    Nothing is proxied to a model we cannot inspect, host and explain. If we cannot run it ourselves, it is not on offer.

  • A catalogue you can follow

    The catalogue is configuration, not code: we add, version and retire models without your applications changing a line, and we tell you when one moves.

Point your application at it

If your code already speaks the OpenAI API, moving to a sovereign endpoint is a base URL and a key. Tell us what you are running and we will size it with you.