LLM Reference

LLM Reference is my go-to for instantly finding and comparing the best AI models and providers tailored to your specific project.

Visit

Published on:

May 29, 2026

Category:

Pricing:

LLM Reference application interface and features

About LLM Reference

LLM Reference is a decision-support directory built for engineers and technology leaders who need to choose the right large language model (LLM) and provider in today's fast-moving AI landscape. It tracks over 1,700 models from more than 130 providers and 235 research labs, with data refreshed weekly to include new releases, verified price changes, and benchmark updates. The core value proposition is simple: stop wasting time hunting through scattered sources and start shipping with confidence. Whether you are building a coding assistant, an agentic workflow, a writing tool, or a research pipeline, LLM Reference gives you a single, trustworthy place to compare models side-by-side, see who offers the cheapest pricing for frontier output, and browse curated editors' picks for specific tasks like coding, agents, writing, research, image generation, and video creation. The site is designed for fast triage. You can quickly identify the right model for your job, determine the most cost-effective provider, and get back to building. With a Pulse feed that highlights what changed this week, including new models, price cuts, and benchmark refreshes, LLM Reference keeps you informed without the noise. It is built by the Data Advantage project and updated daily, making it an essential resource for anyone who needs to stay current with the exploding LLM ecosystem. I personally rely on it to cut through the hype and make real decisions for production deployments.

Features of LLM Reference

This is where I start every single time. The editors' picks are opinionated and curatorial, with personal recommendations for specific tasks. Instead of drowning in a spreadsheet of 1,800 models, you get a shortlist of the best options for coding, agents, writing, research, image generation, and video creation. Each pick comes with a clear rationale, benchmark scores, and a verdict on why it beats the competition. For example, Claude Fable 5 is the current favorite for coding with an 80.3% SWE-bench Pro score, while Claude Opus 4.7 wins for writing with a Chatbot Arena score of 1503. These are informed, expert picks, not just popularity contests.

Live Model and Provider Directory

The directory is the backbone of the product, tracking 1,843 language models from 140 providers and 247 labs. You can search by model name, provider, or task like coding, RAG, agents, long context, vision, classification, or JSON/tool use. Every entry includes pricing, benchmark scores, and provider details. I use this to quickly find the cheapest frontier output, which currently sits at $0.260 per 1M output tokens from Hunyuan HY3 Preview via Tencent Cloud TI Platform. The data is refreshed weekly, so you are always looking at current prices and scores, not stale information from a blog post six months ago.

Pulse Feed for Weekly Changes

The Pulse feed is a dedicated changelog that highlights what actually changed in the model market this week. It tracks three key signals: new models (177 added this week, like DiffusionGemma 26B A4B IT and North Mini Code 1.0), verified price cuts (53 this week), and benchmark refreshes (368 this week across major suites). This is invaluable for staying current without subscribing to ten different newsletters. I check it every Monday morning to see if any of my go-to models have been dethroned or if a price cut makes a new provider more attractive for my current project.

Side-by-Side Model Comparison

The compare tool lets you pit any two models against each other on key dimensions like pricing, benchmark performance, and provider reputation. The cheat sheet on the homepage shows the most-asked comparisons, such as Claude Fable 5 versus Claude Opus 4.8, or GPT-5.5 versus Gemini 3.1 Pro Preview. This feature saves me hours of manual research when I am deciding between two frontier models for a new project. It is clean, fast, and gives you the data you need to make an informed trade-off between performance and cost.

Live Shortlist for Fast Triage

The live shortlist is a dynamic feature that surfaces the best overall model today, the cheapest frontier option, and the freshest update. Right now, it recommends Hunyuan Hy3 Preview as the best overall and cheapest frontier, with DiffusionGemma 26B A4B IT as the freshest update. This is your one-stop shop for a quick decision when you just need to pick a model and get back to building. It is updated in real-time based on the latest data, so you never have to wonder if your choice is still current.

Use Cases of LLM Reference

Selecting the Best Model for a Coding Assistant

When you are building a coding assistant, the choice of model directly impacts code quality and developer productivity. LLM Reference makes this easy by highlighting the top picks for coding, such as Claude Fable 5 with its 80.3% SWE-bench Pro score and 96% SWE-bench Verified score. You can quickly compare it against alternatives like Claude Opus 4.8 or GPT-5.5 to see which one fits your specific use case, whether it is agentic coding, refactoring, or test generation. I used this to pick Claude Sonnet 4.6 for an internal tool because of its strong τ-bench score of 87.5, and it paid off immediately.

Finding the Most Cost-Effective Provider for Production

Cost is a critical factor when scaling LLM usage in production. LLM Reference tracks verified pricing from over 140 providers, so you can find the cheapest frontier output for your task. For example, the current cheapest frontier output is from Hunyuan HY3 Preview at $0.260 per 1M output tokens via Tencent Cloud TI Platform. You can also browse price cuts that are verified weekly, ensuring you are not overpaying for a model that just got cheaper elsewhere. This feature alone has saved my team thousands of dollars by switching to a more cost-effective provider for our summarization pipeline.

Benchmarking a New Model Against the Field

Whenever a new model drops, you need to know how it stacks up against the competition. LLM Reference tracks 1,200 scores across major benchmark suites and refreshes them weekly. You can compare a new model like DiffusionGemma 26B A4B IT against established leaders like Claude Fable 5 or GPT-5.5 on specific benchmarks like coding, reasoning, or long-context tasks. The editors' picks also provide curated context on whether a new model is truly a game-changer or just marketing hype. I used this to quickly dismiss a hyped model that actually underperformed on the benchmarks that mattered for my agent workflow.

Identifying the Right Model for a Research Pipeline

For research pipelines that require deep analysis, summarization, or data extraction, choosing the right model is crucial. LLM Reference provides editors' picks for research, with Claude Fable 5 currently leading due to its GDPval-AA ELO of 1932 and strong performance in finance, trading, and analytics tasks. You can also filter by task like long context or JSON/tool use to find a model that excels at structured output. I recently switched my research pipeline to Claude Opus 4.7 after seeing it top the writing and summarization boards, and the quality of the output improved noticeably.

Frequently Asked Questions

How often is the data in LLM Reference updated?

The data is refreshed weekly to include new model releases, verified price changes, and benchmark updates. The Pulse feed highlights exactly what changed each week, including the number of new models, price cuts, and benchmark refreshes. The site itself is updated daily, so you can trust that the information is current. For example, this week alone saw 177 new models, 53 price cuts, and 368 benchmark refreshes.

The editors' picks are opinionated and curatorial, based on a combination of benchmark performance, real-world task suitability, pricing, and provider reliability. Each pick comes with a detailed rationale, including specific benchmark scores and a clear verdict on why it beats the competition. The goal is to provide a trustworthy starting point for teams who do not have the time to evaluate every model themselves. The picks are updated regularly as new models and benchmark data become available.

Can I compare models from different providers directly?

Yes, the compare tool allows you to pit any two models against each other on key dimensions like pricing, benchmark performance, and provider details. The cheat sheet on the homepage also highlights the most-asked comparisons, such as Claude Fable 5 versus Claude Opus 4.8 or GPT-5.5 versus Gemini 3.1 Pro Preview. This feature is designed to save you time by giving you a clear, side-by-side view of the trade-offs.

The editors' picks cover a wide range of tasks organized by audience. For developers, there are picks for coding, agents, tool use, open weights, long context, and cheap models. For knowledge workers, there are picks for writing, research, summarization, docs Q&A, translation, and data/SQL. For creatives, there are picks for image generation, video, voice (TTS), transcription, music, and image editing. Each category has multiple picks with clear recommendations and rationales.

Similar to LLM Reference

GeoRank

Planning a relocation or long-term stay abroad? Compare places on sunshine, cost, tax, visa and stay duration, then ask AI about your shortlist.

SLABCALC

SlabCalc is a free suite of 34 concrete calculators plus AI tools for crack diagnosis and quote analysis. Instant answers for your concrete project.

BlueHumanizer

Get clear, natural AI text, meaning intact. Free.

Adviserry

Automatically get personalized actions from your YT/Pod/email subs

PrettyScale

Free AI tools to analyze your face attractiveness, find your celebrity look-alike, and calculate your body shape — instant results, no signup.

AuditBadger

SOC 2 and ISO 27001 turned into a clear to-do list. AI prepares the first drafts, you approve every call, and the founders actually answer.

Rizz app - Tchatcheur

AI crafts smooth, ready replies from your chats.

Distro

Distro is an AI Distribution Operator that helps B2B teams publish content, find buyer conversations, engage prospects, and turn social intent into pi