Research

Gwenlake was born out of academic research, and we never stopped: alongside client projects, the team keeps running research projects of its own, and part of it still teaches and publishes. This is where that work surfaces: working prototypes you can try in your browser, and the code we release as open source.

Small models, transformers, reinforcement learning… and games!

A question we meet on every project: when a task is well defined and will run a million times, do you take a large general model and phrase the task in its terms, or do you build a small model around the structure of the task itself?

We took it to two games whose score leaves no room for argument, Tetris and chess.

Read the full write-up

Gwenflow, our agentic AI framework

Gwenflow runs AI agents: give an agent a goal, tools, a memory and a model, and it reasons, acts, observes and goes round again until the goal is met. It is the SDK behind every agentic system we build, and we wrote it ourselves to own every line of code our agents run.

  • Python
  • MIT licence
  • pip install gwenflow
from gwenflow import Agent, ChatMistralfrom gwenflow.tools import Tool def accounts_at_risk(period: str, min_drop: int = 10) -> list[dict]:    """Accounts whose usage fell by at least `min_drop` percent."""    return crm.query(period=period, trend="down", min_drop=min_drop) agent = Agent(    name="Account review",    instructions="Rank the accounts at risk, with the figures.",    llm=ChatMistral(),    tools=[Tool(accounts_at_risk)],) response = agent.run("Which accounts are at risk this quarter?")
  • Any model, one API

    OpenAI, Anthropic, Mistral, Google, DeepSeek, a local Ollama or our own API: swap the provider with one import and the agent code stays the same. Reasoning, streaming and images, audio or PDFs come through the same way.

  • Agents, tools and teams

    An agent runs the loop: call the model, dispatch the tools, read the results, repeat until done. Any Python function becomes a tool. Give an agent a team of specialists and it delegates to them; longer pipelines are DAGs written in code or YAML.

  • Production built in

    Structured output as Pydantic models, a RAG pipeline with document readers and vector stores, skills loaded on demand, and OpenTelemetry tracing out of the box.