Labs
Experiments, benchmarks, simulators and interactive tools to explore AI, agents, robotics, education, productivity and complex systems.
Benchmarks
A saturation tracker for AI-agent benchmarks: what they measure, SOTA, headroom, the human gap and research priority.
An interactive benchmark comparing the academic demandingness of 18 university-entrance exams (PAES, PAU, IB, SAT, Gaokao, Suneung, JEE…) across 7 dimensions, with a world map, radars and ranking.
An interactive benchmark comparing the academic demandingness of Chile's PAES, Spain's PAU and the IB Diploma across 6 cognitive dimensions and 8 subjects, with radar charts.
Simulators
Interactive sandbox: configure an agent on six axes, pick an OWASP LLM Top 10 threat, switch the eighteen controls of the Agentic Control Matrix on and off, and read the coverage — quadrants filled, quadrants left empty, frameworks answered. Deliberately without a residual-risk score.
Calculators
Estimate monthly operating cost, cost per correct outcome, savings, ROI and break-even from workload volume, token prices, tools, retries, human review and rework.
Calculate how many cases you need to detect at least one failure at a chosen confidence level and to estimate its frequency with an explicit margin.
Size review hours, escalations, FTE, cost, sustainable volume and backlog before putting an AI system into production.
Converters
Classify a system across six behavioural axes and its action surfaces, estimate its governance-risk band, surface control signals and find its nearest neighbours among 24 mapped agents.
Crosswalk the 18 Agentic Control Matrix controls to the EU AI Act, ISO 42001, NIST AI RMF, OWASP LLM Top 10 and MITRE ATLAS—with an explicit caveat: mapping is not compliance.
Convert tokens into words, pages, documents, reading and speaking minutes, context-window usage and approximate input cost.
Educational games
A five-round educational game: choose the right governance controls under a limited budget, close each agent's essential gaps and learn from immediate explanations.
Twelve cases to distinguish the base capability, the entity pursuing goals and the infrastructure connecting context, memory, tools and controls.
Investigate AI claims, choose the evidence that actually tests them and spot saturation, contamination, uncertainty and misleading metrics.
A twelve-round game about 18 university-entrance systems: compare demandingness, identify countries and separate cognitive profile from selective pressure.
Experiments
A probabilistic map of where the episodes of the Iliad and the Odyssey might have happened — each location classed as accepted, plausible, speculative or mythical, scored 0–12 against a published rubric, with sources, rival theories and reusable JSON/GeoJSON.
An interactive taxonomy that encodes 24 AI agents — from OpenAI, Anthropic, Google, Microsoft, Amazon and SpaceXAI to the open-source frontier — as vectors across six orthogonal axes plus an action surface, with a filterable table, an A×T×I governance-risk matrix and a machine-readable JSON API.
An interactive map of the 22 occupational groups, comparing AI's theoretical capability against its observed real-world use (as of June 2026), with a category ranking, the most-exposed occupations and the macro figures — rigorously sourced.
Which national team punches above its squad value at the 2026 World Cup. A live efficiency index — points vs €-value — team by team.
An interactive atlas of 100+ world-class waves — wave direction (left / right / both), break type, level and month-by-month average wave height, plotted on a world map.
130+ ski resorts on a world map — skiable km, runs by difficulty, snow quality, monthly snowfall, temperature and sun, plus lift-pass prices.
40+ lift-served MTB / e-bike parks worldwide — trails by difficulty, vertical drop, km, season window and day-pass prices, on a world map.