Models

Start here if you’re new. Otherwise:

  • MODELS — The AI engines themselves. One at a time, directly. You are here.
  • TOOLS — Apps built on top of an AI to do a specific job. Tools
  • ORCHESTRATION & AGENTS — Systems that choose or coordinate which AI answers you. Orchestration
  • TERMS — The words you keep hearing, in plain language. Terms

Have a specific job to do? Go to Tools. Wondering why your AI seems to have gotten worse lately? That’s Orchestration.

Models are the individual “thinking” engines: the models you talk to directly. Everything built on top of them is on Tools.

Each of these models is a different “brain.” They vary in four ways: size and speed, open or closed weights, reasoning on or off, and what they were trained on (and where culture, language and bias live). You should try models trained in other languages, cultures or even time-periods (see below). Even a single model responds differently when you converse with it in different languages. (Anthropic, for example, reports that Claude is more brief, deferential and warm in Arabic than in English.)

The numbering systems and names are confusing. Sakana’s Fugu is not Japan’s Fugaku-LLM. Gemini is not Gemma. A version number means whatever the company wants it to mean.

Benchmark scores are not the whole story. If it is reported by the company it is just marketing, but EpochAI tracks models independently, Arena runs head-to-head leaderboards and MERA scores Russian-language tasks and is a good reminder of why English-only benchmarks mislead. You will need some experience to see how a model performs for YOU.

You will need to log in, even for the free versions (there are few exceptions for brief use). You are logged in to hundreds of other products by these tech giants and logging in also lets you adjust your settings to turn off training. You may or may not want ChatGPT to remember everything about you.

Try the apps, especially voice mode. ChatGPT voice (the blue button, not the microphone) is native and much better than the others at the moment.

They get news from different places: Google has a deal with the Associated Press, Meta uses Fox and CNN, Grok uses X. If you want to know what students are saying about your midterm vs what is happening in DC, you might choose a different model.

All of them have slow, medium and fast models. Hunt for the buttons and drop-down menus. ChatGPT now selects the mode for your prompt, which usually means the cheapest and worst model, so you need to ask for deeper thinking. More on that on Orchestration.

85 of 87 shown

The Big Three

Frontier

  • ChatGPT — OpenAI’s flagship. Broadest feature set in one place: voice, vision, images, code, apps.
  • Claude — Anthropic’s model. Strong on long documents and writing; Cowork and Code for agentic work.
  • Gemini — Google has focused on deep integration with Google products, and is less generally capable except for best-in-class image generation.

Checked 2026-07-24

Other Proprietary Models

Frontier

  • Grok — Designed with extreme personalities and a political bias. Real-time access to X.
  • Copilot — Microsoft’s ChatGPT-based assistant. If your institution provides it in MS Office you may be FERPA and HIPAA secure.
  • Ernie — Baidu’s multimodal frontier model. Ernie X1 is the reasoning version.
  • Sonus — Family of proprietary models: Pro, Air, Mini, and Pro with Reasoning.
  • Amazon Nova — Amazon’s model family, available through Bedrock.
  • MAI-Thinking-1 — Microsoft’s first in-house model. Medium-size, built for math, coding and enterprise use.

Education

  • EduGPT — Built for security, privacy and education use. One of a series of EU models from CivAI.

Privacy

  • Euria — Swiss Infomaniak model. Privacy-focused, renewable powered, recovered energy heats homes.
  • Thaura — Free model from two Syrian engineers. Dedicated to global justice; no training on your data, no military contracts.

Conversation

  • Pi — Focused on emotional intelligence, dialogue and role-playing. Talks but cannot hear.

Search

  • You.com — Started as a privacy-focused search competitor; now focused on specialized API.

Checked 2026-07-24

Open Weight Models

Frontier

  • Inkling — From Mira Murati’s Thinking Machines. Mixture-of-Experts, 1M token context.
  • Kimi — Huge free multimodal Chinese model with a very large context window. Strong on math and coding. Full K3 weights released July 2026 — the largest open-weight model ever.
  • Z.ai (GLM) — Beijing’s GLM series rivals Western proprietary models at lower cost. Usable without logging in. GLM-5.3 (August 2026) runs on the same base model as 5.2 — every gain came from more post-training, which is worth knowing if you think bigger always means better. Weights promised after safety testing.
  • DeepSeek — Strong text and math; built for a fraction of the cost through Multi-head Latent Attention. Free to download.
  • Qwen — Alibaba’s open-weight range with a slider for reasoning level. Guest mode available. In August 2026 Alibaba opened its top-tier model as well: Qwen3.8-Max, 2.4 trillion parameters with 95 billion active and a 1M-token context. The benchmark claims are Alibaba’s own.
  • Meta AI — Meta’s model family; may not require a login.
  • Mistral — French open-source LLM. Fast, more multilingual than the big three, with press-wire news access.
  • MiniMax — Chinese company with open-source reasoning models (M1), video and agent tools.
  • Deep Cogito — Open-weight reasoning models. Current release is Cogito v2.1, a 671B model finetuned from DeepSeek.
  • HuggingChat — Chatbot running on Llama. A good place to see what open source can do. No login required.
  • Nemotron 3.5 Lightning — NVIDIA’s 30B open-weight model, built to be the fast cheap worker an agent calls ten thousand times, not the brain that plans the job. Weights, training data and recipes all released.
  • K2 Horizon — Six models from 0.9B to 375B, released 3 September 2026 by MBZUAI’s Institute of Foundation Models in Abu Dhabi under Apache 2.0, with weights, code and training data. Benchmark claims are the institute’s own.

Reasoning

  • Magistral — Mistral’s reasoning version.
  • MiMo — Xiaomi’s open-source reasoning model.

Checked 2026-07-24

Regional & Culturally-Specific

North America

  • Latimer — Named for Lewis Latimer. Trained further on licensed books, oral histories and sources from Black and Brown communities.

Latin America

  • LatamGPT — Open-weight model from Chile’s CENIA trained on characteristic data from Latin America.
  • Maritaca AI — Brazilian Portuguese chatbot and the Sabia family of fine-tuned models.
  • Rio 3.5 Open 397B — Open-weight Portuguese and English model from IplanRIO, the City of Rio de Janeiro’s IT company, under an MIT license. Built by merging two models and distilling from a stronger one; the model card says the weights first uploaded were the wrong version and a correction is pending.

Middle East

  • Fanar — Culturally and regionally aware Arabic LLM from Qatar’s MCIT and QCRI, fluent in Arabic dialects.
  • Falcon — UAE open-weight model using state space architecture instead of transformers. Handles many dialects.
  • Mistral Saba — 24B Mistral trained on curated Middle East and South Asian data. Arabic and Indian-origin languages.
  • Jais — Trained on the largest Arabic dataset used for an open-source foundational model. Inception rebranded to Inception42; Jais 2 released with Cerebras and MBZUAI.
  • Qalam — Arabic writing and translation tool with plagiarism detection and sentiment analysis.
  • Humain — Saudi model in Arabic and English, working with Qualcomm. Also makes an AI laptop.
  • PersianLLaMA — 13B model fine-tuned on Persian Wikipedia.
  • Karnak Chat — Egyptian government chat assistant built on Karnak, the national Arabic model from the Ministry of Communications’ Applied Innovation Center. Handles Modern Standard Arabic and Egyptian colloquial. Beta; sign-in required.
  • K2 Horizon — Six models from 0.9B to 375B, released 3 September 2026 by MBZUAI’s Institute of Foundation Models in Abu Dhabi under Apache 2.0, with weights, code and training data. Benchmark claims are the institute’s own.

Europe and Russia

  • Soofi S — German open-weight model funded by the Federal Ministry for Economic Affairs. Industrial and enterprise focus.
  • AMALIA — Fully open-source 9B model natively trained in European Portuguese. Launched July 2026.
  • Apertus — Swiss National Supercomputing Centre. 1000 languages, 40% non-English data, fully open development.
  • OpenLLM-Ro — Set of LLaMA fine-tuned Romanian models.
  • Bielik — Open-source family with deep integration of Polish language and cultural context.
  • GigaChat — Open source Russian language model. Useful as a cultural model even where others score higher on MERA.
  • UCCIX — Irish language LLM.
  • Latxa — Family of LLaMA-based Basque models.

Africa

  • Vulavula — Speaks, transcribes and translates low-resource African languages including isiZulu, Sesotho and African French.
  • EthioLLM — Trained on Amharic, Ge’ez, Afan Oromo, Somali, Tigrinya and English.

South Asia

  • Airavata — 7B model optimized for Hindi, from AI4Bharat.
  • Nanda — Hindi-English open-weight model from G42 and MBZUAI. NANDA 87B, built on Llama 3.1 and trained on 65 billion Hindi tokens, released December 2025. No public model page could be found.
  • Sarvam — India’s flagship AI lab: models built for Indian languages with India-hosted inference. A trillion-parameter model is announced but not yet released.
  • BharatGen — India’s government-funded multimodal LLM initiative, fluent in 22+ Indian languages. Anchored by IIT Bombay.

Southeast Asia

  • SEA-Lion — Open source LLM for Southeast Asian languages: Indonesian, Thai, Vietnamese, Filipino, Burmese, Khmer, Malay, Tamil, Lao.
  • Komodo — First LLM built on Indonesian languages including Banjarese, Sundanese, Balinese and Toba Batak.
  • Typhoon — Set of open source models optimized for Thai.

East Asia

  • Doubao — ByteDance model with voice mode; one of the most popular AIs in China.
  • Hy — Tencent’s Hy3 (July 2026, formerly Hunyuan) is an open-weight MoE: 295B total parameters, only 21B active, 256K context, Apache 2.0 licensed.
  • TAIDE — 7B model grounded in Taiwanese culture, language, values and customs.
  • Fugaku-LLM — Japanese-language model built on the RIKEN supercomputer Fugaku. Unrelated to Sakana’s Fugu.
  • HyperCLOVA — Naver’s model, designed to understand Korean in its societal context.

Central Asia

  • Sherkala — Pre-trained on Kazakh and English with some Russian and Turkish.

Checked 2026-07-24

Science & Math

Research

  • Edison Platform — Single environment for scientific R&D. Commercial spinout of FutureHouse with multiple agents.
  • Kosmos — Lab-in-the-loop AI scientist. One run reads ~1,500 papers and executes ~42,000 lines of code, fully auditable.
  • Lila — Scientific AI connected to autonomous laboratories. Trained on 10 trillion tokens of scientific reasoning data.

Earth Science

  • OlmoEarth — Foundation models for Earth observation and remote sensing: mangrove change, forest loss, crop-type maps.

Math

  • WolframAlpha — Combines Wolfram Alpha’s computational power with ChatGPT.
  • AlphaGeometry 2 — Competes at gold-medal IMO level. Not generally available.
  • Numina-Lean-Agent — Uses standard models to do formal math reasoning.
  • Mathstral — Built on Mistral 7B for scientific discovery and mathematical reasoning. Open source.
  • Aristotle — Harmonic’s theorem prover. Takes a problem in English, formalises it, and returns a machine-checked proof in Lean; it can also work inside an existing Lean project. Its IMO gold-medal result is reported in the company’s own preprint. Login required.

Physics

  • Walrus — Trained on real scientific datasets rather than text. Domain is fluids and fluid-like systems.
  • MatterChat — Berkeley Lab bridge model connecting an LLM with physics-based interatomic potentials.

Astronomy

  • AION-1 — Trained on data from astronomical surveys. Polymathic AI collaboration with Cambridge.

Chemistry

  • Ether0 — 24B reasoning model built on Mistral-Small-24B, post-trained for chemistry.

Materials

  • SCIGEN — MIT model steering generation toward materials with quantum structures like Kagome lattices.

Checked 2026-07-24

World Models

A world model generates a space you move through rather than a clip you watch: you steer a camera and the model renders what comes next, keeping objects where you left them. That persistence is the difference from a video generator. Access is the current constraint – only one of these three is open to subscribers today.

Frontier

  • Atlas — World Labs’ omni world model: generates, reconstructs and simulates 3D scenes from text, images, video or 3D data. Early access by request; World Labs says it will power future versions of Marble.
  • Solaris — Runway’s interface world model: it draws an interactive app frame by frame as you use it, with no code underneath. Early access by request.
  • Genie 3 — Google DeepMind’s world model, reached through Project Genie: a text or image prompt becomes an environment you walk through in real time. Requires a Google AI Ultra subscription and a US account.
  • Marble — World Labs’ world model. Generates an explorable 3D scene from text, an image, a video or a rough 3D layout. Its Chisel tool blocks out scene structure in simple shapes first, then applies style separately.

Checked 2026-09-03

Historical LLMs

Historical

  • Mr Chatterbox (Victorian) — Trained from scratch on 28,000 Victorian-era British texts, 1837-1899, from the British Library.
  • TimeCapsuleLLM — Student-built model trained on 1800-1875 London texts to speak Victorian.
  • Talkie — 13B model trained only on pre-1931 public domain text.
  • MonadGPT — Trained on 11,000 texts from 1400-1700 using 17th-century knowledge frameworks.
  • Latin BERT — Contextual language model for Latin.
  • XunziALLM — Uses the formal rules of classical Chinese poetry.
  • CuneiScribe — Reads Akkadian clay tablets and writes new ones. Type English, get cuneiform rendered as a tablet image, and it tells you when it is guessing. Published as TabletCraft (arXiv 2608.02609). Its own Gilgamesh demo mistranslates the famous opening line.

Checked 2026-07-24

Small Models & Edge AI

Small

  • Phi-4 — Microsoft’s small open-weight family, with a reasoning variant. Supersedes Phi-3.5.
  • OpenELM — Apple’s small model family, 270M-3B parameters.
  • Gemma — Google’s open source smaller model, several sizes. Not the same as Gemini.

Checked 2026-07-24

Classifiers & Decision Models

These don’t write. They take text in and return a decision (classify, route, score, pick) with a confidence score, fast and cheap enough to sit inside a workflow. The waiter, not the chef.

Classifiers

  • Jev — TypeSafe AI’s first “System One” model (Sept 2026). No text output: typed decisions with probabilities. Claims 193x faster and 445x cheaper than frontier models on its own tests; early independent tests are small but favorable, and one found its raw confidence scores overconfident.

Checked 2026-09-22

How Models Are Scored

Benchmarks

  • EpochAI — Independent tracking of models and capabilities, with a benchmarks dashboard and a dataset of notable models.
  • Arena — Head-to-head leaderboards, including text-to-image. Renamed from LMArena, formerly Chatbot Arena.
  • MERA — Leaderboard for Russian-language tasks. A good example of why English-only benchmarks mislead.

Tracking

Checked 2026-07-24