Skip to content

Notebook

Notes from the work

We test new models, benchmarks and deployment patterns as they appear, and write about what holds up in real companies.

Latest article

Latest6 min read

Jev: a model that returns a decision, not a sentence

TypeSafe AI released a model on 15 September that writes no text at all. It hands back a chosen option and a probability, costs $0.042 per million input tokens, and charges nothing for output. Here is what survives once the marketing is subtracted.

ModelsCostsClassification
Read

All articles

7 min read

Agent skills: why five beat a hundred

With five skills in the pool, 29.6% of the skills an agent actually uses are the right one; with a hundred, 3.3%. And in August a public skills registry served clones that stole SSH keys. Four rules for a team working with agents.

Read
8 min read

What a language model actually costs to run

From $0.40 to $50 per million tokens in September 2026 list prices, a context window the model only partly uses, and two separate numbers that both get called latency. Figures you can multiply against your own process.

Read
5 min read

Foundation models are learning numbers, not just words

TabPFN predicts on tables without training a model, Revolut trained one on 24 billion events, and Synthefy is raising money for time series. What it means for a company with data but no AI team.

Read
8 min read

Agents in e-commerce: what works today, measured

In the retail domain of a public benchmark, an agent passes 79% of cases on the first attempt and 60% when the same case is run four times. The second number decides whether it can face your customers.

Read
8 min read

Cleora: recommendations without a language model

An open-source graph embedding engine built in Poland. What its authors actually measured, what such a recommendation costs next to a language model, and the places where the method stops working.

Read
8 min read

How you know the agent actually works

An agent passed 69 per cent of cases on the first attempt and 46 per cent when every case had to succeed four times running. The numbers, the judge agreement rates and the cost of measuring systems that never answer the same way twice.

Read

Show us the process that costs your team the most time

Describe it in a few sentences. We’ll tell you whether it can be improved, roughly what that would cost, and whether it needs AI at all.