- Pondhouse Data OG - We know data & AI
- Posts
- Pondhouse Data AI - Tips & Tutorials for Data & AI 59
Pondhouse Data AI - Tips & Tutorials for Data & AI 59
Jev Makes Typed AI Decisions in Milliseconds | GPT-6.1 Sol Nears Astra at One-Fifth the Price | Own Your LLM Inference

Hey there,
This week is about using the right model for the job: TypeSafe's Jev makes fast, typed decisions instead of writing text, our tutorial shows what changes when you run your own inference server, and Pi is a minimal coding agent you can shape to your workflow.
We'll also look at Claude Opus 5.5 and Sonnet 5.5, GPT-6.1 Sol, AI agents doing real science, and a practical tip on why the harness around your coding model can change costs by up to 5×.
Enjoy the read!
Cheers, Andreas & Sascha
In today's edition:
📚 Tutorial of the Week: Own Your Inference: A Hands-On vLLM Workshop
🛠️ Tool Spotlight: Pi: A Minimal, Hackable Coding Agent
📰 Top News: Jev: Typed AI Decisions in Milliseconds
💡 Tip: Benchmark Model and Harness Together
Let's get started!
Tutorial of the week
Own Your Inference: A Hands-On vLLM Workshop
Most AI agents rent their intelligence: every call goes to someone else's API, priced per token and subject to their rate limits. Akamai's open "Agents That Own Their Inference" workshop shows what changes when you run the model yourself, with Jupyter notebooks that work against your own vLLM server on a dedicated GPU.
See where a request spends its time: prefill vs. decode, time to first token vs. time per output token, and why decode is memory-bound.
Size and speed up your model: count parameters, precision, VRAM, bandwidth and KV cache, then switch from BF16 to FP8 and add speculative decoding, measuring whether each change actually helps.
Push the server under load: drive traffic until batching, queueing and KV-cache pressure show up in vLLM's metrics, then tune serving settings and keep only what improves throughput or latency.
Follow a structured path: ten numbered modules end with a capstone that puts an agent on the endpoint you tuned. Modules 1–4 only need a vLLM endpoint, so you can start before your cluster is ready.
Avoid lock-in: the stack is open source (vLLM, Kubernetes, Qwen models) and runs on any Kubernetes cluster with one NVIDIA GPU. A self-run GPU node is billed by the hour.
If you already call LLM APIs and want to understand the layer your agents actually run on, this is a well-structured place to start. Basic Python and container knowledge is enough.
Tool of the week
Pi: A Minimal, Hackable Coding Agent for the Terminal
Most coding agents come with a fixed workflow and a preferred model provider. Pi, an MIT-licensed project from Earendil Works, takes the opposite approach: it starts minimal, and you ask the agent itself to build the prompt templates, skills and extensions you need, or install ready-made Pi packages.
Use any provider: one unified API for OpenAI, Anthropic, Google and others.
Automate or embed it: run Pi interactively, script it in print, JSON or RPC mode, or build it into your own apps with the TypeScript SDK.
Reuse the building blocks: the repository also ships the agent runtime (tool calling, state management) and the terminal UI library.
Keep long sessions lean: instead of trimming large tool logs, Pi saves the full output to disk and lets the model search it when needed. Pi's dev notes, as reported by AlphaSignal, cite 26–35% less context and up to 88% lower processing costs across 19 sessions.
Sandbox it yourself: Pi has no built-in permission system and runs with your user's permissions. For untrusted work, use a Gondolin micro-VM, Docker or an OpenShell sandbox.
With more than 110,000 GitHub stars, Pi is worth a look if you want a coding agent that adapts to your workflow instead of the other way around.
Top News of the week
TypeSafe's Jev Makes Typed AI Decisions in Milliseconds

TypeSafe AI, founded by former OpenAI researcher Diogo Almeida, has opened early access to Jev, its first "System One" model. Jev doesn't write text. You define questions up front, such as "Is this ticket urgent?" or "Which team should handle it?", and Jev returns an answer for each, with a probability for every allowed option, in a single parallel pass that always matches your schema.
The hype comes from speed and price: TypeSafe reports 70–500 ms per call and $0.042 per million input tokens, with output free. On its own workflow benchmarks, it claims about two orders of magnitude lower latency and cost than frontier LLMs, which the company itself calls the high end of real-world gains. The intended uses are "smart if-statements" in ordinary code: classify, route, score, extract, and judge or guardrail LLM output.
Jev is not an LLM replacement. It can't write text, code or explanations, and a valid output isn't necessarily a correct one. An independent Carnegie Mellon study, JEV-as-a-Judge, found it within three points of GPT-6 on verdicts that can be read off the text, at 0.36% of the fee, but behind on math, code and logic. Accepting Jev's confident verdicts and escalating the rest to a reasoning model was slightly more accurate than GPT-6 alone, at 41% of the cost.
Jev is less a rival to LLMs than a new building block next to them.
Also in the news
Anthropic's Claude 5.5 Models Get Faster and Cheaper
Anthropic released Claude Opus 5.5 and Sonnet 5.5. Opus 5.5 costs about 40% less than Opus 5 on typical workloads and generates output 30% faster, while Sonnet 5.5 keeps its price but needs fewer tokens per task. Check the migration notes before upgrading: Opus 5.5 has four breaking changes, including thinking that can no longer be disabled and no forced tool use.
GPT-6.1 Sol Nears GPT-6 Astra at One-Fifth the Price
One week after halving GPT-6 Sol and Luna prices, OpenAI launched GPT-6.1 Sol at DevDay. OpenAI says it nearly matches its flagship GPT-6 Astra on agentic coding, computer use and professional work at one-fifth of Astra's token price. According to TechCrunch, citing the Wall Street Journal, OpenAI dropped a planned GPT-6.1 Astra over safety concerns.
AMD Buys World Labs, Makes Fei-Fei Li Its Chief Scientist
AMD has acquired World Labs, whose Atlas world model was our top story last week. Fei-Fei Li becomes AMD's Executive Vice President and Chief Scientist, working directly with CEO Lisa Su, and the World Labs team continues its work inside AMD. Both say they want to build an open AI ecosystem spanning hardware, software and open models.
Claude Agents Find an Unknown Enzyme System With CRISPR-Like Repeats
In the first result from its new life sciences lab, Anthropic reports that about 950 Claude agents searched a DNA database for 21 hours and spotted a CRISPR-like repeat array next to an unusual reverse transcriptase in bacteriophages. Humans only wrote the prompt and ran the lab work. What the system, called ART, actually does is still unknown, but CRISPR pioneer Feng Zhang called it "genuinely intriguing."
Google's ScientistTwo Beats 86 of 107 Published ML Methods
Google Cloud AI Research introduced ScientistTwo, a multi-agent system that takes a research problem and autonomously runs baselines, tests new ideas, and writes the paper and code. On 107 problems from ICLR, ICML and NeurIPS papers, it beat the published state-of-the-art method in 86 cases, by 25.2% on average, according to the authors. Its papers also out-scored accepted ones, but only in automated AI reviews, not human peer review.
Tip of the week
Benchmark Coding Agents as Model–Harness Pairs

Switching to a cheaper model is not the only way to cut coding-agent costs. The harness around it, the software that runs the loop, manages context and calls tools, can matter just as much.
The HarnessTax study tested 21 model–harness pairs: seven models across Claude Code, Codex CLI and Pi (our tool of the week), on SWE-bench Lite and Terminal-Bench 2.0. The same model could cost up to 5× more depending on the harness, with little change in how many tasks it solved.
The practical lesson is to treat model and harness as one choice:
Test them together: run each model with more than one harness on a small set of your own representative tasks.
Change one variable at a time: keep the model fixed while comparing harnesses, then keep the harness fixed while comparing models.
Track cost per solved task: pass rates hide the difference. For example, if one setup spends $40 and another $12 to solve the same 8 tasks, both score the same, but one is more than three times cheaper.
Re-run after changes: repeat the comparison whenever you switch models or update your harness.
As the study puts it, your Claude models may not need Claude Code. Measure before you assume.
We hope you liked our newsletter and you stay tuned for the next edition. If you need help with your AI tasks and implementations - let us know. We are happy to help

