Client work from Tesseract Academy and Tesseract Advisory & Consulting
Proven in production

Case studies

Production AI, research and corporate AI training across retail, fintech, Web3, insurance, healthtech, sport, education, HR and the public sector — what the client needed, what we built, and what changed. Numbers appear exactly as reported on each project; where an outcome was qualitative, we say so.

89%
of churning customers predicted, with precision up to 90% — insurance
≈£1m/yr
what a device insurer can save with claims forecasts up to 50% better than experts
70%
valid simulation environments generated with LLMs — with the Alan Turing Institute
5,012 drafts30-day pilot · every reply human-reviewed
Generative AI · Multi-agent systems · Customer support · Evaluation

Production AI support agent with earned autonomy

A global career-services group running three consumer brands handled a high volume of repetitive billing, refund, cancellation and account email across 19 locales. Off-the-shelf LLM support bots were rejected as unsafe: a wrong refund promise or a fabricated account fact reaches a real customer.

What we did
  • Built a multi-agent system (OpenAI Agents SDK, FastAPI, Postgres) embedded in the client's live support inbox: a triage agent routes to four specialists, each grounded in live account, payment and refund-eligibility data — no canned templates.
  • Set the quality bar empirically: statistical analysis of 6,734 real human support conversations produced the scoring rubric and three significant drivers of reply quality, wired directly into prompt compilation.
  • Built a four-track evaluation platform (LLM-as-judge G-Eval scorers, deterministic compliance validators, CI gates) plus a three-tier guardrail pipeline for safety, policy and tone.
  • Ran six adversarial red-team audits against live production drafts, each closing a specific failure class (fabricated account claims, internal-instruction leakage, over-promised financial actions).
  • Designed staged autonomy: write-actions ship feature-flagged and off, earning autonomy per action only after audit evidence. Money never moves automatically.
In pilot on real traffic the agent drafted 5,012 replies in 30 days, every one human-reviewed before send. Against the human-baseline rubric its drafts outscore the average human reply on 79% of conversations (composite 0.78 vs 0.62). Fabricated financial-action claims and internal-instruction leaks measured at zero in the latest audit, down from recurring. Delivered in five months, with 1,200+ automated tests.
≈£1m/yrpotential savings in stock & supply chain
Insurance · Supply chain

Claims forecasting for a world-leading electronics insurer

Replacement stock for 1,000+ device models at different lifecycle stages: overbuy and money is wasted; understock and claims can't be serviced on time.

What we did
  • Benchmarked exponential smoothing, Kalman filters, Prophet and ARIMA against Tesseract's custom AI forecaster.
  • The forecaster learns pattern variations across all devices and detects which trajectory a new device follows.
  • Handled concept drift — launch-to-adoption pattern shifts — by recognising similar curves from other devices.
  • Benchmarked final output against human-expert predictions.
Up to 50% better forecasts than expert predictions in many cases — worth close to £1 million a year, with savings accumulating year on year.
89%of churn predicted · up to 90% precision
Insurance

Customer churn prediction under privacy constraints

A leading London device-insurance provider needed to predict active and passive churn — with customer demographics off-limits under privacy rules.

What we did
  • Compared classification (risk scores) against survival modelling (risk over time).
  • Identified and ranked the factors that put customers at higher risk.
  • Built an ML pipeline predicting who churns and when, so retention outreach could be prioritised.
About 89% of churning customers predicted with precision up to 90% — roughly 4 of every 5 future churners can be reached before they leave.
The Alan Turing Institute
Research · Cybersecurity

LLM-generated cyber-attack simulations, with the Alan Turing Institute

As a Strategic Partner of the Institute, we co-delivered a study using LLMs to generate cybersecurity simulation environments and enhance reinforcement-learning threat simulation.

What we did
  • Tested template-based vs example-based LLM generation of YAML network configurations — example-based won.
  • Benchmarked three RL attack agents — including a Cyber Kill Chain agent simulating multi-stage attacks — against classical approaches including PPO.
70% of generated environments were valid and simulation-ready — research-grade methods that feed directly into client engagements. See our research & publications.
IEEE-publishedclinical validation vs polysomnography
Medical statistics · HealthTech

Clinical validation of a contactless sleep monitor

Polysomnography is the clinical gold standard for sleep staging, but it needs a lab and EEG electrodes — impractical for long-term home monitoring. A healthtech company's radar-based contactless monitor needed rigorous, publishable validation.

What we did
  • Benchmarked the radar-based monitor and its sleep-analysis algorithm epoch-by-epoch against clinical polysomnography.
  • Repeated the evaluation on an independent validation dataset (n=24) to test robustness and generalisability.
  • Published the methodology and results at IEEE EMBC 2020 (peer-reviewed; Dr Kampakis co-author).
Peer-reviewed sleep-stage recall of 75.0% (deep), 74.8% (REM), 59.9% (light) and 57.1% (wake), holding on the independent dataset — published at IEEE EMBC 2020. The paper · our research.
GOV.WALESpublished national research · cited in the Senedd
Public sector · Statistics

Land valuation research for the Welsh Government

The Welsh Government commissioned independent research into the feasibility of land-value tax models for Wales — a statistical question with national policy consequences.

What we did
  • Tested five land-valuation methodologies across 99% of Welsh geography.
  • Presented findings to Welsh Government officials.
Published on GOV.WALES (March 2026) as "Testing land valuation methods" and cited in Senedd committee proceedings — commissioned, delivered, on the public record.
51,355 triplesnational skills ontology · zero SHACL violations
Ontology · GovTech

Ontology engineering for the public sector

Government data is deeply relational but published document-shaped — schemas unpublished, crosswalks buried, relationships lost in spreadsheets. Making it machine-reasonable is what makes AI usable in the public sector.

What we did
  • Delivered an open-source AI ontology extension tool for the National Digital Twin Programme (Department for Business and Trade, 2025; Apache-2.0).
  • Run an open research programme: rebuilt the Skills England occupational maps as a 51,355-triple national ontology validating with zero SHACL violations (independent demonstration, all 1,269 standards placed), plus published crosswalks between the UK's government ontology standards.
A commissioned government delivery plus an open, citable research programme — explore it at gov.tesseract.academy/research.
Peer-reviewedtokenomics audit framework · tested on Terra/Luna & Ethereum 2.0
Tokenomics · Web3

A published method for auditing token economies

Token economies fail in public and at scale — many projects launch on ad-hoc token issuance with no economic stress-testing. Investors and foundations need a way to audit a token design before it goes live.

What we did
  • Developed the Tokenomics Audit Checklist — peer-reviewed in the Journal of The British Blockchain Association (2023) and applied to a DeFi project, Terra/Luna and Ethereum 2.0.
  • Built TokenLab, an open-source agent-based simulator that stress-tests token economies against speculator behaviour before launch (arXiv, 2024).
A repeatable, peer-reviewed audit framework — four journal papers from "Why do we need Tokenomics?" (2018) to the audit checklist (2023) — plus open-source simulation tooling. The checklist paper · TokenLab · our research.
JBBA 2022peer-reviewed audit of a live stablecoin
Web3 · Stablecoins

Auditing a live stablecoin’s token economy

Stablecoins live or die by their economic design — collateral mechanics, incentive loops and failure modes. A live stablecoin project needed its token economy independently audited.

What we did
  • Audited the project’s live token economy — assessing viability and suggesting improvements, as an independent view of the design.
  • Set out general principles for auditing a token economy, demonstrated on the live project.
  • Published the method and lessons as a peer-reviewed case study in the Journal of The British Blockchain Association (2022).
A peer-reviewed audit of a live stablecoin project — followed by the Tokenomics Audit Checklist (2023). The paper · our research.
3 engagementstoken-design engagements · peer-reviewed (JBBA 2018)
Web3 · Token design

Token design in practice: three real engagements

Three projects, three different economic problems — token designs that had to hold up after launch, not just in the whitepaper.

What we did
  • Took on three real token-design engagements, each with a distinct economic challenge.
  • Documented how each challenge was analysed and solved, and published all three as a peer-reviewed case-study paper (JBBA, 2018).
  • Alongside the case for the discipline itself: “Why do we need Tokenomics?” (JBBA, 2018).
Three real token-design engagements, peer-reviewed as case studies in the Journal of The British Blockchain Association. The paper · our research.
Premier Leagueinjury-prediction research with Tottenham Hotspur FC
Sports analytics · Machine learning

Predicting football injuries for an elite club

Player injuries can cost elite clubs millions per season, and the data that could predict them — training load, GPS traces, injury histories — is messy, sparse and noisy. Exactly the kind of operational data most organisations struggle to model.

What we did
  • Doctoral research at UCL in collaboration with Tottenham Hotspur FC: machine learning on training-load, GPS and injury data to predict football injuries and recovery.
  • Benchmarked ML methods for predicting a player's recovery time after undiagnosed injury (ECML/PKDD MLSA workshop, 2013).
  • Extended the toolkit across sport: T20 cricket match-outcome models built from 500+ features that beat a gambling-industry benchmark (2015), and ML player-valuation modelling (arXiv, 2022).
A defended UCL PhD plus published sports-analytics research — including models that outperformed a gambling-industry benchmark — the origin of our expertise in ML on messy operational data. The thesis · our research.
500+ trainedInnovate UK BridgeAI · creative industries
Education · National programme

Training 500+ professionals for Innovate UK’s BridgeAI programme

The UK’s creative industries have high growth potential but low AI adoption — the gap Innovate UK’s BridgeAI programme exists to close.

What we did
  • Awarded the “Creative Industries AI Training and Support” contract by Innovate UK BridgeAI.
  • Delivered a fully funded AI upskilling programme for UK creative professionals — online modules, group calls and applied projects.
  • The work was featured at the BridgeAI Annual Showcase.
500+ professionals trained through the programme — a national, publicly funded delivery. The announcement.
50+ trainedat Vodafone in Egypt · in-house delivery
Education · Corporate training

In-house AI training for organisations like Vodafone and British Land

Generic, off-the-shelf AI training rarely survives contact with a real organisation. Enterprises bring us in to train their own people, on their own context.

What we did
  • Trained 50+ people for Vodafone in Egypt.
  • Delivered AI training for British Land.
  • The same delivery model now runs as the Agentic Flywheel — training that becomes workflows.
In-house delivery for Vodafone and British Land — see how we work with companies.
70+ toolsopen-source AI-native ontology engine
Ontology · Open source

Open Ontologies: an AI-native ontology engine, in the open

Knowledge graphs are how organisations make their data machine-reasonable — but the tooling has been heavyweight, JVM-bound and never built for AI-assisted workflows.

What we did
  • Built Open Ontologies in the open (MIT), led by our partner Fabio Rovai — a Rust MCP server exposing 70+ tools for building, validating, querying and reasoning over RDF/OWL ontologies.
  • Native OWL2-DL reasoning, SHACL validation and SPARQL, with a desktop Studio — no JVM required.
  • Applied case studies in the open research programme at gov.tesseract.academy/research are built with it.
An open engine anyone can run — MIT-licensed and actively developed. Open Ontologies on GitHub.
LLM + RAGrated flawless by a human expert
Fintech

LLMs and RAG for portfolio reporting

The client's engine for turning portfolio performance into natural language was fragile and error-prone.

What we did
  • Refined a custom LLM integrated with a RAG system encoding portfolio knowledge.
  • Trained it via data augmentation: a second LLM generated thousands of examples from a small expert-provided seed set.
Results judged flawless by a human expert — and the client now holds unique IP they are using to secure further funding.
New fraudsurfaced beyond known cases
Financial services

An AI fraud-detection engine, piloted in the real world

A startup set out to offer AI fraud detection as a service — piloted with a financial services company managing a large intermediary network with a known fraud problem.

What we did
  • Started with an AI Roadmap to weigh the ways forward and their trade-offs.
  • Built an ensemble engine combining outlier detection, ML and statistical techniques.
  • Enlarged the dataset via statistical augmentation; iterated the pilot against ground truth.
The engine caught known fraud and surfaced new cases — and the client raised a further investment round on the pilot's success.
<30 minto analyse any dataset
HR · People analytics

Interpretable AI for organisational culture data

A culture-analytics startup had huge volumes of sensitive, confounded questionnaire data — and needed certainty that a factor like gender or age truly drives a behaviour before reporting it to management.

What we did
  • Built a proprietary interpretable-AI algorithm fitting a statistical model to every question.
  • Summarised those models with ML, then distilled each demographic factor's true effect.
Any dataset analysed in under 30 minutes with all significant factors distilled — the client has grown strongly since and is raising a major round.
In productionautomated pricing decisions
Retail

Dynamic pricing for a UK retailer

Real-time pricing that responds to demand, competitor prices and inventory levels as shopping shifts to digital channels.

What we did
  • Started with an AI roadmap: data assets, quality, candidate approaches, and the benefits, risks and costs of each.
  • Built and trained a pricing model on the client's own data over six months, then put it into production.
The model now makes automated pricing decisions in production, improving margins and reducing inventory costs.

Want results like these?

Start with the 2-minute readiness assessment, or see how the Agentic Flywheel turns training into shipped workflows.

Start the assessment → For companies