Map LLM Latency, Cost, and Accuracy.
A desktop client for testing, grading, and benchmarking prompts across major LLMs.
A Local Tool. Not a Cloud Subscription.
- Bring your own API keys.
- Run concurrent evals across multiple models.
- Version-control your prompt iterations.
- Compare model outputs side-by-side.
- Export leaderboards and download charts as PNGs.
- Everything writes to a local SQLite file.
- One-time purchase, no subscriptions.
Model Benchmarking.
Run your golden dataset against multiple models simultaneously. Review the batch outputs side-by-side, assign pass/fail grades, and isolate the exact model that handles your edge cases.
Version Control for Your Prompts.
Keep a clean history of every iteration. Fork a prompt to test a new variable, track the exact changes, and easily revert to past versions without losing your context.
Request-Level Debugging.
Chat interfaces hide the details. Inspect the raw API responses, latency stats, and exact token consumption for every single request.
Global Model Configuration.
Connect your APIs and set baseline parameters once. No more manual toggling for every test run.
Provider & Key Management
Bring your own keys for OpenAI, Anthropic, Mistral, Gemini, and xAI. Define your baseline testing stack globally so your go-to models are automatically active the second you create a new prompt.
Global Inference Defaults
Set baseline configurations for max completion tokens, temperature, top_p, and model-specific parameters like reasoning effort. Defaults apply to all runs unless explicitly overridden at the prompt level.
Find the Right LLM for Your Dataset.
Run concurrent evaluations locally across OpenAI, Anthropic, Gemini, and xAI.
Try the app for free. $29 for a permanent license.
Learn about our permanent licensing and early-adopter pricing:
View License Details →