Model Library · US

214 models

Available models with live pricing, context windows, and status.

More filters
Max input priceAnyEGP / 1M
Max output priceAnyEGP / 1M

Context window

Providers

214 models matching
Sovereign · Hosted in Egypt

2 sovereign models

Served from our own GPUs inside Egypt — prompts and completions never leave the country. Built for data-residency and compliance-sensitive workloads.

About sovereign hosting →
Sovereign · Hosted in Egypt
Alibaba
ChatOpen weights

Qwen3 is the latest generation of large language models in Qwen series, offering a comprehensive suite of dense and mixture-of-experts (MoE) models. Built upon extensive training, Qwen3 delivers groundbreaking advancements in reasoning, instruction-following, agent capabilities, and multilingual support, with the following key features:

Context

4K

In EGP / 1M

5.00

Out EGP / 1M

10.00

Qwen/Qwen3-0.6B
Sovereign · Hosted in Egypt
Alibaba
ChatOpen weights

Following the widespread community adoption of the Qwen3.5 and Qwen3.6 series, we are pleased to introduce Qwen3.8, the most capable generation in the Qwen open-model family to date.

Context

131K

In EGP / 1M

14.00

Out EGP / 1M

105.00

SovereignEG/Qwen3.8-27B-FP8

Global catalog

212
EmbeddingOpen weights

We present a sentence transformation model that generates semantically similar sentences. Our model is based on the Sentence-Transformers architecture and was trained on a large dataset of sentence pairs. We evaluate the effectiveness of our model by measuring its ability to generate similar sentences that are close to the original sentence in meaning.

Context

512

In EGP / 1M

0.28

Out EGP / 1M

Free

EmbeddingOpen weights

We present a sentence transformation model that achieves state-of-the-art results on various NLP tasks without requiring task-specific architectures or fine-tuning. Our approach leverages contrastive learning and utilizes a variety of datasets to learn robust sentence representations. We evaluate our model on several benchmarks and demonstrate its effectiveness in various applications such as text classification, sentiment analysis, named entity recognition, and question answering.

Context

512

In EGP / 1M

0.28

Out EGP / 1M

Free

EmbeddingOpen weights

A sentence transformation model that has been trained on a wide range of datasets, including but not limited to S2ORC, WikiAnwers, PAQ, Stack Exchange, and Yahoo! Answers. Our model can be used for various NLP tasks such as clustering, sentiment analysis, and question answering.

Context

512

In EGP / 1M

0.28

Out EGP / 1M

Free

ChatOpen weights

Context

33K

In EGP / 1M

5.60

Out EGP / 1M

5.60

qwen-2-1.5b-instruct
EmbeddingOpen weights

BGE embedding is a general Embedding Model. It is pre-trained using retromae and trained on large-scale pair data using contrastive learning. Note that the goal of pre-training is to reconstruct the text, and the pre-trained model cannot be used for similarity calculation directly, it needs to be fine-tuned

Context

512

In EGP / 1M

0.28

Out EGP / 1M

Free

BAAI
EmbeddingOpen weights

A LLM-based embedding model with in-context learning capabilities that achieves SOTA performance on BEIR and AIR-Bench. It leverages few-shot examples to enhance task performance.

Context

8K

In EGP / 1M

0.56

Out EGP / 1M

Free

EmbeddingOpen weights

BGE embedding is a general Embedding Model. It is pre-trained using retromae and trained on large-scale pair data using contrastive learning. Note that the goal of pre-training is to reconstruct the text, and the pre-trained model cannot be used for similarity calculation directly, it needs to be fine-tuned

Context

512

In EGP / 1M

0.56

Out EGP / 1M

Free

bge-m3

Live
BAAI
EmbeddingOpen weights

BGE-M3 is a versatile text embedding model that supports multi-functionality, multi-linguality, and multi-granularity, allowing it to perform dense retrieval, multi-vector retrieval, and sparse retrieval in over 100 languages and with input sizes up to 8192 tokens. The model can be used in a retrieval pipeline with hybrid retrieval and re-ranking to achieve higher accuracy and stronger generalization capabilities. BGE-M3 has shown state-of-the-art performance on several benchmarks, including MKQA, MLDR, and NarritiveQA, and can be used as a drop-in replacement for other embedding models like DPR and BGE-v1.5.

Context

8K

In EGP / 1M

0.56

Out EGP / 1M

Free

BAAI
EmbeddingOpen weights

BGE-M3 is a multilingual text embedding model developed by BAAI, distinguished by its Multi-Linguality (supporting 100+ languages), Multi-Functionality (unified dense, multi-vector, and sparse retrieval), and Multi-Granularity (handling inputs from short queries to 8192-token documents). It achieves state-of-the-art retrieval performance across diverse benchmarks while maintaining a single model for multiple retrieval modes.

Context

512

In EGP / 1M

0.56

Out EGP / 1M

Free

EmbeddingOpen weights

BGE-M3 is a multilingual text embedding model developed by BAAI, distinguished by its Multi-Linguality (supporting 100+ languages), Multi-Functionality (unified dense, multi-vector, and sparse retrieval), and Multi-Granularity (handling inputs from short queries to long documents). It achieves state-of-the-art retrieval performance across diverse benchmarks while maintaining a single model for multiple retrieval modes. This endpoint serves the model's full 8192-token context.

Context

8K

In EGP / 1M

0.56

Out EGP / 1M

Free

Anthropic
ChatProprietary

Claude Fable 5.1 improves on Claude Fable 5 across the board, with the biggest gains in agentic coding, long-running agentic workflows, and knowledge work: long code refactors, front-end and visual

Context

1M

In EGP / 1M

560.07

Out EGP / 1M

2800.32

claude-fable-5.1
Anthropic
ChatProprietary

Claude Opus 4.5 is Anthropic’s frontier reasoning model optimized for complex software engineering, agentic workflows, and long-horizon computer use. It offers strong multimodal capabilities, competitive performance across real-world coding and

Context

200K

In EGP / 1M

280.03

Out EGP / 1M

1400.16

claude-opus-4.5
Anthropic
ChatProprietary

Opus 4.6 is Anthropic’s strongest model for coding and long-running professional tasks. It is built for agents that operate across entire workflows rather than single prompts, making it especially effective

Context

1M

In EGP / 1M

280.03

Out EGP / 1M

1400.16

claude-opus-4.6
Anthropic
ChatProprietary

Claude Opus 5 is Anthropic’s flagship model for demanding reasoning, coding, and long-horizon agentic work. It is particularly strong at end-to-end software tasks, code review and bug finding, visual analysis

Context

1M

In EGP / 1M

560.07

Out EGP / 1M

2800.32

claude-opus-5
Anthropic
ChatProprietary

Claude Sonnet 4.5 is Anthropic’s most advanced Sonnet model to date, optimized for real-world agents and coding workflows. It delivers state-of-the-art performance on coding benchmarks such as SWE-bench Verified, with

Context

1M

In EGP / 1M

168.02

Out EGP / 1M

840.10

claude-sonnet-4.5
Anthropic
ChatProprietary

Claude Fable 5 is Anthropic's next generation of intelligence for the hardest knowledge work and coding problems. It works independently for longer than any prior generally available Claude model: run it in an agent harness and it can work for days at a time, planning across stages, delegating to sub-agents, and checking its own work.

Context

1M

In EGP / 1M

560.07

Out EGP / 1M

2800.32

Anthropic
ChatProprietary

The next generation of Anthropic's fastest and most cost-effective model, optimal for use cases where speed and affordability matter.

Context

200K

In EGP / 1M

56.01

Out EGP / 1M

280.03

Anthropic
ChatProprietary

Anthropic's most capable production model yet, advancing performance across coding, enterprise workflows, and long-running agentic tasks.

Context

1M

In EGP / 1M

280.03

Out EGP / 1M

1400.16

Anthropic
ChatProprietary

Claude Opus 4.8 is our most intelligent Opus model and the best generally available model for coding and agents, with deeper reasoning for enterprise workflows.

Context

1M

In EGP / 1M

560.07

Out EGP / 1M

2800.32

Anthropic
ChatProprietary

Claude Sonnet 4.6 delivers frontier intelligence at scale—built for coding, agents, and enterprise workflows.

Context

1M

In EGP / 1M

168.02

Out EGP / 1M

840.10

Anthropic
ChatProprietary

Claude Sonnet 5 is Anthropic's most capable Sonnet model yet, built for coding, agents, and professional work at scale. It brings near-Opus intelligence to the model teams run at scale every day, with the same balance of capability, cost, and speed teams already rely on Sonnet for.

Context

1M

In EGP / 1M

112.01

Out EGP / 1M

560.07

SBERT
EmbeddingOpen weights

The CLIP model maps text and images to a shared vector space, enabling various applications such as image search, zero-shot image classification, and image clustering. The model can be used easily after installation, and its performance is demonstrated through zero-shot ImageNet validation set accuracy scores. Multilingual versions of the model are also available for 50+ languages.

Context

77

In EGP / 1M

0.28

Out EGP / 1M

Free

EmbeddingOpen weights

This model is a multilingual version of the OpenAI CLIP-ViT-B32 model, which maps text and images to a common dense vector space. It includes a text embedding model that works for 50+ languages and an image encoder from CLIP. The model was trained using Multilingual Knowledge Distillation, where a multilingual DistilBERT model was trained as a student model to align the vector space of the original CLIP image encoder across many languages.

Context

512

In EGP / 1M

0.28

Out EGP / 1M

Free

ChatOpen weights

DeepSeek R1 Distill Llama 70B is a distilled large language model based on Llama-3.3-70B-Instruct, using outputs from DeepSeek R1. The model combines advanced distillation techniques to achieve high performance across

Context

8K

In EGP / 1M

44.81

Out EGP / 1M

44.81

deepseek-r1-distill-llama-70b
DeepSeek
ChatOpen weights

The DeepSeek R1 model has undergone a minor version upgrade, with the current version being DeepSeek-R1-0528.

Context

164K

In EGP / 1M

28.00

Out EGP / 1M

120.41

DeepSeek
ChatOpen weights

DeepSeek-V3, a strong Mixture-of-Experts (MoE) language model with 671B total parameters with 37B activated for each token. To achieve efficient inference and cost-effective training, DeepSeek-V3 adopts Multi-head Latent Attention (MLA) and DeepSeekMoE architectures, which were thoroughly validated in DeepSeek-V2.

Context

164K

In EGP / 1M

17.92

Out EGP / 1M

49.85

DeepSeek
ChatOpen weights

DeepSeek-V3.1 is post-trained on the top of DeepSeek-V3.1-Base, which is built upon the original V3 base checkpoint through a two-phase long context extension approach, following the methodology outlined in the original DeepSeek-V3 report. We have expanded our dataset by collecting additional long documents and substantially extending both training phases. The 32K extension phase has been increased 10-fold to 630B tokens, while the 128K extension phase has been extended by 3.3x to 209B tokens. Additionally, DeepSeek-V3.1 is trained using the UE8M0 FP8 scale data format to ensure compatibility with microscaling data formats.

Context

164K

In EGP / 1M

14.00

Out EGP / 1M

53.21

ChatOpen weights

DeepSeek-V3.1 Terminus is an update to DeepSeek V3.1 that maintains the model's original capabilities while addressing issues reported by users, including language consistency and agent capabilities, further optimizing the model's performance in coding and search agents. It is a large hybrid reasoning model (671B parameters, 37B active) that supports both thinking and non-thinking modes. It extends the DeepSeek-V3 base with a two-phase long-context training process. Users can control the reasoning behaviour with the reasoning enabled boolean. Learn more in our docs The model improves tool use, code generation, and reasoning efficiency, achieving performance comparable to DeepSeek-R1 on difficult benchmarks while responding more quickly. It supports structured tool calling, code agents, and search agents, making it suitable for research, coding, and agentic workflows.

Context

131K

In EGP / 1M

15.12

Out EGP / 1M

56.01

DeepSeek
ChatOpen weights

DeepSeek-V3.2 is a large language model designed to harmonize high computational efficiency with strong reasoning and agentic tool-use performance. It introduces DeepSeek Sparse Attention (DSA), a fine-grained sparse attention mechanism that reduces training and inference cost while preserving quality in long-context scenarios. A scalable reinforcement learning post-training framework further improves reasoning, with reported performance in the GPT-5 class, and the model has demonstrated gold-medal results on the 2025 IMO and IOI. V3.2 also uses a large-scale agentic task synthesis pipeline to better integrate reasoning into tool-use settings, boosting compliance and generalization in interactive environments.

Context

164K

In EGP / 1M

14.56

Out EGP / 1M

21.28

DeepSeek
ChatOpen weights

DeepSeek V4 Flash is an efficiency-focused MoE model with 284B total parameters (13B active) and a 1M-token context window. It's tuned for fast inference and high-throughput use cases while still holding up on reasoning and coding tasks.

Context

1M

In EGP / 1M

5.04

Out EGP / 1M

10.08

DeepSeek
ChatOpen weights

DeepSeek V4 Pro is an MoE model with 1.6T total parameters (49B active) and a 1M-token context window. It's built for advanced reasoning, coding, and long-running agent tasks, and performs well on knowledge, math, and software engineering benchmarks.

Context

1M

In EGP / 1M

72.81

Out EGP / 1M

145.62

DeepSeek V3, a 685B-parameter, mixture-of-experts model, is the latest iteration of the flagship chat model family from the DeepSeek team. It succeeds the DeepSeek V3 model and performs really well

Context

164K

In EGP / 1M

16.24

Out EGP / 1M

63.85

deepseek-chat-v3-0324
Chat

DeepSeek-V3.1 is a large hybrid reasoning model (671B parameters, 37B active) that supports both thinking and non-thinking modes via prompt templates. It extends the DeepSeek-V3 base with a two-phase long-context

Context

164K

In EGP / 1M

14.00

Out EGP / 1M

53.21

deepseek-chat-v3.1
ChatOpen weights

DeepSeek V4 Flash 0731 is a sparse mixture-of-experts model from DeepSeek, with 13B active parameters out of 284B total. This re-post-trained revision is suited for coding, reasoning, and agent workflows.

Context

1M

In EGP / 1M

3.36

Out EGP / 1M

10.08

deepseek-v4-flash-0731
ChatOpen weights

DeepSeek V4 Flash Vision Exp is an experimental vision-enabled version of DeepSeek V4 Flash 0731 from DeepSeek, adding image understanding while matching the base model on text capabilities including agents,

Context

1M

In EGP / 1M

24.64

Out EGP / 1M

73.93

deepseek-v4-flash-vision-exp
ChatOpen weights

DeepSeek V4 Pro 0813 is a large-scale mixture-of-experts model from DeepSeek. This is the GA release of DeepSeek V4 Pro.

Context

1M

In EGP / 1M

72.81

Out EGP / 1M

145.62

deepseek-v4-pro-0813
DeepSeek
Chat

DeepSeek R1 is here: Performance on par with OpenAI o1, but open-sourced and with fully open reasoning tokens. It's 671B parameters in size, with 37B active in an inference pass

Context

64K

In EGP / 1M

39.20

Out EGP / 1M

140.02

deepseek-r1
intfloat
EmbeddingOpen weights

Text Embeddings by Weakly-Supervised Contrastive Pre-training. Model has 24 layers and 1024 out dim.

Context

512

In EGP / 1M

0.28

Out EGP / 1M

Free

intfloat
EmbeddingOpen weights

Text Embeddings by Weakly-Supervised Contrastive Pre-training. Model has 24 layers and 1024 out dim.

Context

512

In EGP / 1M

0.56

Out EGP / 1M

Free

EmbeddingOpen weights

EmbeddingGemma is a 300M parameter multilingual open embedding model from Google DeepMind, designed for efficient deployment even on low-resource devices, producing high-quality text vector representations for tasks such as search, classification, clustering, and semantic similarity.

Context

2K

In EGP / 1M

0.11

Out EGP / 1M

Free

ChatProprietary

Gemini 2.5 Flash is Google's latest thinking model, designed to tackle increasingly complex problems. It's capable of reasoning through their thoughts before responding, resulting in enhanced performance and improved accuracy. Gemini 2.5 Flash: best for balancing reasoning and speed.

Context

1M

In EGP / 1M

16.80

Out EGP / 1M

140.02

Google
ChatProprietary

Gemini 2.5 Pro is Google's the most advanced thinking model, designed to tackle increasingly complex problems. Gemini 2.5 Pro leads common benchmarks by meaningful margins and showcases strong reasoning and code capabilities. Gemini 2.5 models are thinking models, capable of reasoning through their thoughts before responding, resulting in enhanced performance and improved accuracy. The Gemini 2.5 Pro model is now available on DeepInfra.

Context

1M

In EGP / 1M

70.01

Out EGP / 1M

560.07

ChatProprietary

Nano Banana Pro (Gemini 3 Pro Image) is designed to tackle the most challenging image generation by incorporating state-of-the-art reasoning capabilities. It is the best model for complex and multi-turn image generation and editing.

Context

66K

In EGP / 1M

112.01

Out EGP / 1M

672.08

ChatProprietary

Bring any idea to life with state-of-the-art reasoning to help you learn, build, and plan anything. Best for high-volume tasks that need efficiency and intelligence.

Context

1M

In EGP / 1M

14.00

Out EGP / 1M

84.01

Google
ChatProprietary

Bring any idea to life with state-of-the-art reasoning to help you learn, build, and plan anything. Best for complex tasks and bringing creative concepts to life.

Context

1M

In EGP / 1M

112.01

Out EGP / 1M

672.08

ChatProprietary

Gemini 3.5 Flash delivers near-Pro intelligence at Flash-tier cost and speed: Pro-level coding proficiency, parallel agentic execution, all at a much lower price.

Context

1M

In EGP / 1M

84.01

Out EGP / 1M

504.06

Google
ChatOpen weights

Context

131K

In EGP / 1M

1.12

Out EGP / 1M

5.60

gemma-4-e4b-it
ChatOpen weights

Gemma 2 27B by Google is an open model built from the same research and technology used to create the Gemini models. Gemma models are well-suited for a variety of

Context

8K

In EGP / 1M

36.40

Out EGP / 1M

36.40

gemma-2-27b-it
Google
ChatOpen weights

Gemma 3 introduces multimodality, supporting vision-language input and text outputs. It handles context windows up to 128k tokens, understands over 140 languages, and offers improved math, reasoning, and chat capabilities, including structured outputs and function calling. Gemma 3-12B is Google's latest open source model, successor to Gemma 2

Context

131K

In EGP / 1M

2.80

Out EGP / 1M

8.40

Google
ChatOpen weights

Gemma 3 introduces multimodality, supporting vision-language input and text outputs. It handles context windows up to 128k tokens, understands over 140 languages, and offers improved math, reasoning, and chat capabilities, including structured outputs and function calling. Gemma 3 27B is Google's latest open source model, successor to Gemma 2

Context

131K

In EGP / 1M

4.48

Out EGP / 1M

8.96

Google
ChatOpen weights

Gemma 3 introduces multimodality, supporting vision-language input and text outputs. It handles context windows up to 128k tokens, understands over 140 languages, and offers improved math, reasoning, and chat capabilities, including structured outputs and function calling. Gemma 3-12B is Google's latest open source model, successor to Gemma 2

Context

131K

In EGP / 1M

2.80

Out EGP / 1M

5.60

ChatOpen weights

Efficient, MoE variant of Gemma 4. Gemma is a family of open models built by Google DeepMind. Gemma 4 models are multimodal, handling text and image input and generating text output.

Context

262K

In EGP / 1M

3.92

Out EGP / 1M

19.04

Google
ChatOpen weights

Gemma is a family of open models built by Google DeepMind. Gemma 4 models are multimodal, handling text and image input and generating text output.

Context

262K

In EGP / 1M

7.28

Out EGP / 1M

21.28

ChatOpen weights

Gemma is a family of open models built by Google DeepMind. Gemma 4 models are multimodal, handling text and image input and generating text output.

Context

262K

In EGP / 1M

5.04

Out EGP / 1M

19.04

ChatProprietary

Ultra speed version of gemma-4-31B-it

Context

131K

In EGP / 1M

15.12

Out EGP / 1M

42.56

zai-org
ChatOpen weights

GLM-4.5V is a vision-language foundation model for multimodal agent applications. Built on a Mixture-of-Experts (MoE) architecture with 106B parameters and 12B activated parameters, it achieves state-of-the-art results in video understanding,

Context

66K

In EGP / 1M

33.60

Out EGP / 1M

100.81

glm-4.5v
zai-org
ChatOpen weights

Compared with GLM-4.5, GLM-4.6 brings several key improvements: Longer context window: The context window has been expanded from 128K to 200K tokens, enabling the model to handle more complex agentic tasks. Superior coding performance: The model achieves higher scores on code benchmarks and demonstrates better real-world performance in applications such as Claude Code、Cline、Roo Code and Kilo Code, including improvements in generating visually polished front-end pages. Advanced reasoning: GLM-4.6 shows a clear improvement in reasoning performance and supports tool use during inference, leading to stronger overall capability. More capable agents: GLM-4.6 exhibits stronger performance in tool using and search-based agents, and integrates more effectively within agent frameworks. Refined writing: Better aligns with human preferences in style and readability, and performs more naturally in role-playing scenarios.

Context

203K

In EGP / 1M

28.00

Out EGP / 1M

112.01

zai-org
ChatOpen weights

GLM-4.7 is a state-of-the-art, multilingual Mixture-of-Experts (MoE) language model designed for complex reasoning, agentic coding, and tool use. Building on its predecessor GLM-4.6, it delivers significant improvements across key benchmarks, including multilingual SWE-bench, Terminal Bench, and reasoning-heavy evaluations like HLE. The model features advanced "Interleaved Thinking" and new "Preserved Thinking" modes, allowing it to reason before actions and maintain consistency across long, multi-turn tasks. With 358 billion parameters, GLM-4.7 excels in generating clean code, modern UI elements, and sophisticated reasoning outputs.

Context

203K

In EGP / 1M

22.40

Out EGP / 1M

98.01

zai-org
ChatOpen weights

GLM-4.7-Flash is a 30B-A3B MoE model. As the strongest model in the 30B class, GLM-4.7-Flash offers a new option for lightweight deployment that balances performance and efficiency.

Context

131K

In EGP / 1M

3.39

Out EGP / 1M

22.40

glm-5

Live
zai-org
ChatOpen weights

GLM-5 is an advanced, open-source large language model designed for developers tackling the toughest challenges. It excels at long-context reasoning, multi-step tool orchestration, and complex systems engineering, making it the ideal choice for powering sophisticated agents and applications that require high-level cognitive tasks.

Context

198K

In EGP / 1M

33.60

Out EGP / 1M

107.53

zai-org
ChatOpen weights

GLM-5.1 is Z-AI's next-generation flagship model for agentic engineering, with significantly stronger coding capabilities than its predecessor. It achieves state-of-the-art performance on SWE-Bench Pro and leads GLM-5 by a wide margin on NL2Repo (repo generation) and Terminal-Bench 2.0 (real-world terminal tasks).

Context

203K

In EGP / 1M

58.81

Out EGP / 1M

196.02

zai-org
ChatOpen weights

GLM-5.2 is Z-AI's latest flagship model for long-horizon tasks. It marks a substantial leap in long-horizon task capability over its predecessor GLM-5.1 and, for the first time, delivers that capability on a **solid 1M-token context**.

Context

1M

In EGP / 1M

42.00

Out EGP / 1M

134.42

zai-org
ChatOpen weights

GLM-5.3-Flash is a native multimodal model from Z.ai. It is suited for efficient coding and long-horizon agent tasks. Its hybrid sparse and linear attention architecture maintains accurate long-context behavior while reducing compute overhead.

Context

1M

In EGP / 1M

8.40

Out EGP / 1M

28.00

Gemini 3.6 Flash is a high-efficiency model from Google for coding, agentic workflows, and web and app development. It is designed to produce polished outputs with fewer unnecessary edits and

Context

1M

In EGP / 1M

42.00

Out EGP / 1M

210.02

gemini-3.6-flash
ChatProprietary

Gemini 3.7 Flash is a multimodal model from Google for fast agentic workflows, coding, and complex multi-step reasoning. It is designed for tasks that require responsive performance and reliable multi-step

Context

1M

In EGP / 1M

42.00

Out EGP / 1M

210.02

gemini-3.7-flash

Gemini 3.8 Flash is Google's most intelligent Flash model with significant gains from 3.7 Flash across software engineering, agentic tasks, and multi-step reasoning.

Context

1M

In EGP / 1M

42.00

Out EGP / 1M

210.02

gemini-3.8-flash
OpenAI
Chat

GPT-3.5 Turbo is OpenAI's fastest model. It can understand and generate natural language or code, and is optimized for chat and traditional completion tasks. Training data up to Sep 2021.

Context

16K

In EGP / 1M

28.00

Out EGP / 1M

84.01

Chat

This model offers four times the context length of gpt-3.5-turbo, allowing it to support approximately 20 pages of text in a single request at a higher cost. Training data: up

Context

16K

In EGP / 1M

168.02

Out EGP / 1M

224.03

This model is a variant of GPT-3.5 Turbo tuned for instructional prompts and omitting chat-related optimizations. Training data: up to Sep 2021.

Context

4K

In EGP / 1M

84.01

Out EGP / 1M

112.01

gpt-4

Live
OpenAI
Chat

OpenAI's flagship model, GPT-4 is a large-scale multimodal language model capable of solving difficult problems with greater accuracy than previous models due to its broader general knowledge and advanced reasoning

Context

8K

In EGP / 1M

1680.19

Out EGP / 1M

3360.39

OpenAI
Chat

OpenAI's flagship model, GPT-4 is a large-scale multimodal language model capable of solving difficult problems with greater accuracy than previous models due to its broader general knowledge and advanced reasoning

Context

8K

In EGP / 1M

1680.19

Out EGP / 1M

3360.39

OpenAI
Chat

GPT-4.1 is a flagship large language model optimized for advanced instruction following, real-world software engineering, and long-context reasoning. It supports a 1 million token context window and outperforms GPT-4o and

Context

1M

In EGP / 1M

112.01

Out EGP / 1M

448.05

OpenAI
Chat

GPT-4.1 Mini is a mid-sized model delivering performance competitive with GPT-4o at substantially lower latency and cost. It retains a 1 million token context window and scores 45.1% on hard

Context

1M

In EGP / 1M

22.40

Out EGP / 1M

89.61

OpenAI
Chat

For tasks that demand low latency, GPT‑4.1 nano is the fastest and cheapest model in the GPT-4.1 series. It delivers exceptional performance at a small size with its 1 million

Context

1M

In EGP / 1M

5.60

Out EGP / 1M

22.40

gpt-4o

Live
OpenAI
Chat

GPT-4o ("o" for "omni") is OpenAI's latest AI model, supporting both text and image inputs with text outputs. It maintains the intelligence level of GPT-4 Turbo while being twice as

Context

128K

In EGP / 1M

140.02

Out EGP / 1M

560.07

Chat

GPT-4o ("o" for "omni") is OpenAI's latest AI model, supporting both text and image inputs with text outputs. It maintains the intelligence level of GPT-4 Turbo while being twice as

Context

128K

In EGP / 1M

280.03

Out EGP / 1M

840.10

Chat

The 2024-08-06 version of GPT-4o offers improved performance in structured outputs, with the ability to supply a JSON schema in the respone_format. Read more here. GPT-4o ("o" for "omni") is

Context

128K

In EGP / 1M

140.02

Out EGP / 1M

560.07

Chat

The 2024-11-20 version of GPT-4o offers a leveled-up creative writing ability with more natural, engaging, and tailored writing to improve relevance & readability. It’s also better at working with uploaded

Context

128K

In EGP / 1M

140.02

Out EGP / 1M

560.07

OpenAI
Chat

GPT-4o mini is OpenAI's newest model after GPT-4 Omni, supporting both text and image inputs with text outputs. As their most advanced small model, it is many multiples more affordable

Context

128K

In EGP / 1M

8.40

Out EGP / 1M

33.60

GPT-4o mini is OpenAI's newest model after GPT-4 Omni, supporting both text and image inputs with text outputs. As their most advanced small model, it is many multiples more affordable

Context

128K

In EGP / 1M

8.40

Out EGP / 1M

33.60

GPT-4o mini Search Preview is a specialized model for web search in Chat Completions. It is trained to understand and execute web search queries.

Context

128K

In EGP / 1M

8.58

Out EGP / 1M

34.32

Chat

GPT-4o Search Previewis a specialized model for web search in Chat Completions. It is trained to understand and execute web search queries.

Context

128K

In EGP / 1M

143.00

Out EGP / 1M

572.00

gpt-5

Live
OpenAI
Chat

GPT-5 is OpenAI’s most advanced model, offering major improvements in reasoning, code quality, and user experience. It is optimized for complex tasks that require step-by-step reasoning, instruction following, and accuracy

Context

400K

In EGP / 1M

70.01

Out EGP / 1M

560.07

OpenAI
Chat

GPT-5 Mini is a compact version of GPT-5, designed to handle lighter-weight reasoning tasks. It provides the same instruction-following and safety-tuning benefits as GPT-5, but with reduced latency and cost

Context

400K

In EGP / 1M

14.00

Out EGP / 1M

112.01

OpenAI
Chat

GPT-5-Nano is the smallest and fastest variant in the GPT-5 system, optimized for developer tools, rapid interactions, and ultra-low latency environments. While limited in reasoning depth compared to its larger

Context

400K

In EGP / 1M

2.80

Out EGP / 1M

22.40

OpenAI
Chat

GPT-5.1 is the latest frontier-grade model in the GPT-5 series, offering stronger general-purpose reasoning, improved instruction adherence, and a more natural conversational style compared to GPT-5. It uses adaptive reasoning

Context

400K

In EGP / 1M

70.01

Out EGP / 1M

560.07

OpenAI
Chat

GPT-5.1-Codex is a specialized version of GPT-5.1 optimized for software engineering and coding workflows. It is designed for both interactive development sessions and long, independent execution of complex engineering tasks

Context

400K

In EGP / 1M

70.01

Out EGP / 1M

560.07

Chat

GPT-5.1-Codex-Max is OpenAI’s latest agentic coding model, designed for long-running, high-context software development tasks. It is based on an updated version of the 5.1 reasoning stack and trained on agentic

Context

400K

In EGP / 1M

70.01

Out EGP / 1M

560.07

Chat

GPT-5.1-Codex-Mini is a smaller and faster version of GPT-5.1-Codex

Context

400K

In EGP / 1M

14.00

Out EGP / 1M

112.01

OpenAI
Chat

GPT-5.2 is the latest frontier-grade model in the GPT-5 series, offering stronger agentic and long context perfomance compared to GPT-5.1. It uses adaptive reasoning to allocate computation dynamically, responding quickly

Context

400K

In EGP / 1M

98.01

Out EGP / 1M

784.09

OpenAI
Chat

GPT-5.2-Codex is an upgraded version of GPT-5.1-Codex optimized for software engineering and coding workflows. It is designed for both interactive development sessions and long, independent execution of complex engineering tasks

Context

400K

In EGP / 1M

98.01

Out EGP / 1M

784.09

OpenAI
Chat

GPT-5.2 Pro is OpenAI’s most advanced model, offering major improvements in agentic coding and long context performance over GPT-5 Pro. It is optimized for complex tasks that require step-by-step reasoning,

Context

400K

In EGP / 1M

1176.14

Out EGP / 1M

9409.09

OpenAI
Chat

GPT-5.3-Codex is OpenAI’s most advanced agentic coding model, combining the frontier software engineering performance of GPT-5.2-Codex with the broader reasoning and professional knowledge capabilities of GPT-5.2. It achieves state-of-the-art results

Context

400K

In EGP / 1M

98.01

Out EGP / 1M

784.09

OpenAI
Chat

GPT-5.4 is OpenAI’s latest frontier model, unifying the Codex and GPT lines into a single system. It features a 1M+ token context window (922K input, 128K output) with support for

Context

1M

In EGP / 1M

140.02

Out EGP / 1M

840.10

OpenAI
Chat

GPT-5.4 mini brings the core capabilities of GPT-5.4 to a faster, more efficient model optimized for high-throughput workloads. It supports text and image inputs with strong performance across reasoning, coding,

Context

400K

In EGP / 1M

42.00

Out EGP / 1M

252.03

OpenAI
Chat

GPT-5.4 nano is the most lightweight and cost-efficient variant of the GPT-5.4 family, optimized for speed-critical and high-volume tasks. It supports text and image inputs and is designed for low-latency

Context

400K

In EGP / 1M

11.20

Out EGP / 1M

70.01

OpenAI
Chat

GPT-5.5 is OpenAI’s frontier model designed for complex professional workloads, building on GPT-5.4 with stronger reasoning, higher reliability, and improved token efficiency on hard tasks. It features a 1M+ token

Context

1M

In EGP / 1M

280.03

Out EGP / 1M

1680.19

OpenAI
Chat

GPT-5.6 Luna is a fast, cost-efficient model in OpenAI's GPT-5.6 series. It is suited for high-volume, latency-sensitive tasks such as chat, classification, and lightweight agentic workflows, providing capable reasoning for

Context

1M

In EGP / 1M

22.40

Out EGP / 1M

134.42

OpenAI
Chat

GPT-5.6 Sol is the flagship model in OpenAI's GPT-5.6 series. It is suited for complex reasoning, coding, and agentic workflows, and is particularly strong at command-line and multi-step coding tasks

Context

1M

In EGP / 1M

112.01

Out EGP / 1M

560.07

OpenAI
Chat

GPT-5.6 Terra is a balanced model in OpenAI's GPT-5.6 series, positioned between the flagship Sol tier and the cost-efficient Luna tier. It is suited for everyday coding, reasoning, and agentic

Context

1M

In EGP / 1M

112.01

Out EGP / 1M

672.08

OpenAI
Chat

GPT-6 Astra is OpenAI's flagship model for demanding end-to-end work. It is suited for advanced analysis, software engineering, deep research, scientific work, and document creation, with particular strengths in long-horizon

Context

1M

In EGP / 1M

280.03

Out EGP / 1M

1400.16

OpenAI
ChatOpen weights

gpt-oss-120b is an open-weight, 117B-parameter Mixture-of-Experts (MoE) language model from OpenAI designed for high-reasoning, agentic, and general-purpose production use cases. The model supports configurable reasoning depth, full chain-of-thought access, and native tool use, including function calling, browsing, and structured output generation.

Context

131K

In EGP / 1M

2.07

Out EGP / 1M

9.52

ChatOpen weights

Context

131K

In EGP / 1M

8.40

Out EGP / 1M

33.60

OpenAI
ChatOpen weights

gpt-oss-20b is an open-weight 21B parameter model released by OpenAI under the Apache 2.0 license. It uses a Mixture-of-Experts (MoE) architecture with 3.6B active parameters per forward pass, optimized for lower-latency inference. The model is trained in OpenAI’s Harmony response format and supports reasoning level configuration, fine-tuning, and agentic capabilities including function calling, tool use, and structured outputs.

Context

131K

In EGP / 1M

1.68

Out EGP / 1M

7.84

ibm-granite
ChatOpen weights

Granite-4.2-30B is the flagship reasoning model in the Granite 4.2 family. It delivers the strongest performance across reasoning-intensive tasks by leveraging built-in <think>...</think> chain-of-thought. It supports flexible thinking modes — full thinking (default), non-thinking, and low-effort — allowing users to balance depth vs. latency on a per-query basis.

Context

131K

In EGP / 1M

8.96

Out EGP / 1M

36.40

ibm-granite
ChatOpen weights

Granite-4.2-3B is the compact reasoning model in the Granite 4.2 family. Despite its small parameter count, it delivers strong performance on reasoning-intensive tasks by leveraging built-in <think>...</think> chain-of-thought. It supports flexible thinking modes — full thinking (default), non-thinking, and low-effort — allowing users to balance depth vs. latency on a per-query basis.

Context

131K

In EGP / 1M

1.68

Out EGP / 1M

6.72

ibm-granite
ChatOpen weights

Granite-4.2-8B is the mid-size reasoning model in the Granite 4.2 family. It delivers strong performance on reasoning-intensive tasks by leveraging built-in <think>...</think> chain-of-thought. It supports flexible thinking modes — full thinking (default), non-thinking, and low-effort — allowing users to balance depth vs. latency on a per-query basis.

Context

131K

In EGP / 1M

3.36

Out EGP / 1M

14.00

thenlper
EmbeddingOpen weights

The GTE models are trained by Alibaba DAMO Academy. They are mainly based on the BERT framework and currently offer three different sizes of models, including GTE-large, GTE-base, and GTE-small. The GTE models are trained on a large-scale corpus of relevance text pairs, covering a wide range of domains and scenarios. This enables the GTE models to be applied to various downstream tasks of text embeddings, including information retrieval, semantic textual similarity, text reranking, etc.

Context

512

In EGP / 1M

0.28

Out EGP / 1M

Free

thenlper
EmbeddingOpen weights

The GTE models are trained by Alibaba DAMO Academy. They are mainly based on the BERT framework and currently offer three different sizes of models, including GTE-large, GTE-base, and GTE-small. The GTE models are trained on a large-scale corpus of relevance text pairs, covering a wide range of domains and scenarios. This enables the GTE models to be applied to various downstream tasks of text embeddings, including information retrieval, semantic textual similarity, text reranking, etc.

Context

512

In EGP / 1M

0.56

Out EGP / 1M

Free

NousResearch
ChatOpen weights

Hermes 3 is a generalist language model with many improvements over Hermes 2, including advanced agentic capabilities, much better roleplaying, reasoning, multi-turn conversation, long context coherence, and improvements across the board.

Context

131K

In EGP / 1M

39.20

Out EGP / 1M

39.20

hy3

Live
tencent
ChatOpen weights

Hy3 is a 295B-parameter Mixture-of-Experts (MoE) model with 21B active parameters and 3.8B MTP layer parameters, developed by the Tencent Hy Team. Following the Hy3 Preview launch in late April, we gathered feedback from 50+ products and scaled up post-training with higher quality data. Today, we introduce Hy3, which outperforms similar-size models and rivals flagship open-source models with 2-5x parameters. It also shows significant gains in utility across various products and productivity tasks.

Context

262K

In EGP / 1M

7.84

Out EGP / 1M

32.48

thinkingmachines
ChatOpen weights

Inkling is an open-weight multimodal mixture-of-experts model from Thinking Machines Lab, with 41B active parameters out of 975B total. It is designed for general-purpose reasoning, coding, agentic and tool-use systems,

Context

524K

In EGP / 1M

53.21

Out EGP / 1M

226.83

inkling
moonshotai
ChatOpen weights

Kimi K2.5 is an open-source, native multimodal agentic model built through continual pretraining on approximately 15 trillion mixed visual and text tokens atop Kimi-K2-Base. It seamlessly integrates vision and language understanding with advanced agentic capabilities, instant and thinking modes, as well as conversational and agentic paradigms.

Context

262K

In EGP / 1M

25.20

Out EGP / 1M

126.01

moonshotai
ChatOpen weights

Kimi K2.6 is an open-source, native multimodal agentic model that advances practical capabilities in long-horizon coding, coding-driven design, proactive autonomous execution, and swarm-based task orchestration.

Context

262K

In EGP / 1M

42.00

Out EGP / 1M

196.02

moonshotai
ChatOpen weights

Kimi K2.7 Code is a coding-focused agentic model built upon Kimi K2.6. With substantial improvements on real-world long-horizon coding tasks, it strengthens end-to-end task completion across complex software engineering workflows while improving token efficiency, reducing thinking-token usage by approximately 30% compared with Kimi K2.6.

Context

262K

In EGP / 1M

38.08

Out EGP / 1M

190.42

inclusionai
ChatOpen weights

*Ling-3.0-flash* is a *124B-parameter Mixture-of-Experts (MoE) model*, with approximately *5.1B parameters activated per token*. The model is designed with *token efficiency and production-scale agentic inference* as key priorities, enabling developers

Context

131K

In EGP / 1M

3.36

Out EGP / 1M

10.08

ling-3.0-flash
inclusionAI
ChatOpen weights

Ling-3.0-flash-Fin is the first finance-enhanced model in the Ant Ling family. Developed by Ant Group with leading financial institutions and domain experts, it extends Ling-3.0-flash through continued training on high-quality financial data.

Context

262K

In EGP / 1M

3.36

Out EGP / 1M

10.08

inclusionAI
Chat

The multimodal version built on Ling-3.0-flash — 124B total / ~5.5B active per token, with native text, image, and video understanding. It’s mainly designed for multimodal agentic workflows, long-context understanding, and multi-step reasoning.

Context

131K

In EGP / 1M

3.36

Out EGP / 1M

10.08

ChatOpen weights

Llama 3.3-70B Turbo is a highly optimized version of the Llama 3.3-70B model, utilizing FP8 quantization to deliver significantly faster inference speeds with a minor trade-off in accuracy. The model is designed to be helpful, safe, and flexible, with a focus on responsible deployment and mitigating potential risks such as bias, toxicity, and misinformation. It achieves state-of-the-art performance on various benchmarks, including conversational tasks, language translation, and text generation.

Context

131K

In EGP / 1M

5.60

Out EGP / 1M

17.92

The Llama 4 collection of models are natively multimodal AI models that enable text and multimodal experiences. These models leverage a mixture-of-experts architecture to offer industry-leading performance in text and image understanding. Llama 4 Maverick, a 17 billion parameter model with 128 experts

Context

1M

In EGP / 1M

11.20

Out EGP / 1M

44.81

ChatOpen weights

The Llama 4 collection of models are natively multimodal AI models that enable text and multimodal experiences. These models leverage a mixture-of-experts architecture to offer industry-leading performance in text and image understanding. Llama 4 Scout, a 17 billion parameter model with 16 experts

Context

328K

In EGP / 1M

5.60

Out EGP / 1M

16.80

ChatOpen weights

Llama Guard 4 is a natively multimodal safety classifier with 12 billion parameters trained jointly on text and multiple images. Llama Guard 4 is a dense architecture pruned from the Llama 4 Scout pre-trained model and fine-tuned for content safety classification. Similar to previous versions, it can be used to classify content in both LLM inputs (prompt classification) and in LLM responses (response classification). It itself acts as an LLM: it generates text in its output that indicates whether a given prompt or response is safe or unsafe, and if unsafe, it also lists the content categories violated.

Context

164K

In EGP / 1M

10.08

Out EGP / 1M

10.08

EmbeddingOpen weights

The llama-nemotron-embed-vl-1b-v2 is a high-performance multimodal embedding model designed to transform text queries and document images into dense vector representations for advanced retrieval systems. It excels at understanding complex visual content like charts, tables, and infographics.

Context

10K

In EGP / 1M

0.56

Out EGP / 1M

Free

ChatOpen weights

Llama 3.2 1B is a 1-billion-parameter language model focused on efficiently performing natural language tasks, such as summarization, dialogue, and multilingual text analysis. Its smaller size allows it to operate

Context

60K

In EGP / 1M

1.51

Out EGP / 1M

11.26

llama-3.2-1b-instruct
ChatOpen weights

Llama 3.2 3B is a 3-billion-parameter multilingual large language model, optimized for advanced natural language processing tasks like dialogue generation, reasoning, and summarization. Designed with the latest transformer architecture, it

Context

131K

In EGP / 1M

2.80

Out EGP / 1M

18.48

llama-3.2-3b-instruct
ChatOpen weights

Meta developed and released the Meta Llama 3.1 family of large language models (LLMs), a collection of pretrained and instruction tuned generative text models in 8B, 70B and 405B sizes

Context

131K

In EGP / 1M

22.40

Out EGP / 1M

22.40

ChatOpen weights

Meta developed and released the Meta Llama 3.1 family of large language models (LLMs), a collection of pretrained and instruction tuned generative text models in 8B, 70B and 405B sizes

Context

131K

In EGP / 1M

1.12

Out EGP / 1M

2.24

XiaomiMiMo
ChatOpen weights

MiMo-V2.5 is a native omnimodal model with strong agentic capabilities, supporting text, image, video, and audio understanding within a unified architecture. Built upon the MiMo-V2-Flash backbone and extended with dedicated vision and audio encoders, it delivers robust performance across multimodal perception, long-context reasoning, and agentic workflows.

Context

262K

In EGP / 1M

22.40

Out EGP / 1M

112.01

XiaomiMiMo
ChatOpen weights

MiMo-V2.5-Pro is an open-source Mixture-of-Experts (MoE) language model with 1.02T total parameters and 42B active parameters. It utilizes the hybrid attention architecture and 3-layers Multi-Token Prediction (MTP) introduced in MiMo-V2-Flash.

Context

1M

In EGP / 1M

56.01

Out EGP / 1M

168.02

MiniMaxAI
ChatOpen weights

MiniMax-M2 is a compact, high-efficiency large language model optimized for end-to-end coding and agentic workflows. With 10 billion activated parameters (230 billion total), it delivers near-frontier intelligence across general reasoning,

Context

205K

In EGP / 1M

14.28

Out EGP / 1M

57.13

minimax-m2
MiniMaxAI
ChatOpen weights

MiniMax-M2.7 is MiniMax's first model deeply participating in its own evolution. M2.7 is capable of building complex agent harnesses and completing highly elaborate productivity tasks, leveraging Agent Teams, complex Skills, and dynamic tool search.

Context

205K

In EGP / 1M

16.80

Out EGP / 1M

67.21

MiniMaxAI
ChatProprietary

Speed-optimized MiniMax-M2.7

Context

197K

In EGP / 1M

21.28

Out EGP / 1M

95.21

MiniMaxAI
ChatOpen weights

MiniMax-M3 is a native multimodal model with 1M context. It has ~428B parameters and ~23B activated parameters.

Context

524K

In EGP / 1M

15.68

Out EGP / 1M

61.61

ChatOpen weights

12B model trained jointly by Mistral AI and NVIDIA, it significantly outperforms existing models smaller or similar in size.

Context

131K

In EGP / 1M

1.06

Out EGP / 1M

1.68

ChatOpen weights

Mistral Small 3 is a 24B-parameter language model optimized for low-latency performance across common AI tasks. Released under the Apache 2.0 license, it features both pre-trained and instruction-tuned versions designed for efficient local deployment. The model achieves 81% accuracy on the MMLU benchmark and performs competitively with larger models like Llama 3.3 70B and Qwen 32B, while operating at three times the speed on equivalent hardware.

Context

33K

In EGP / 1M

2.80

Out EGP / 1M

4.48

ChatOpen weights

Mistral-Small-3.2-24B-Instruct is a drop-in upgrade over the 3.1 release, with markedly better instruction following, roughly half the infinite-generation errors, and a more robust function-calling interface—while otherwise matching or slightly improving on all previous text and vision benchmarks.

Context

128K

In EGP / 1M

4.20

Out EGP / 1M

11.20

Chat

Kimi K2 Instruct is a large-scale Mixture-of-Experts (MoE) language model developed by Moonshot AI, featuring 1 trillion total parameters with 32 billion active per forward pass. It is optimized for

Context

131K

In EGP / 1M

31.92

Out EGP / 1M

128.81

kimi-k2
Chat

Kimi K2 0905 is the September update of Kimi K2 0711. It is a large-scale Mixture-of-Experts (MoE) language model developed by Moonshot AI, featuring 1 trillion total parameters with 32

Context

262K

In EGP / 1M

33.60

Out EGP / 1M

140.02

kimi-k2-0905
moonshotai
ChatOpen weights

Kimi K3 is a 2.8T parameter open-weight multimodal reasoning model from Moonshot AI. It is suited for complex coding, knowledge work, and long-horizon agentic workflows, and is particularly strong at

Context

1M

In EGP / 1M

159.62

Out EGP / 1M

798.09

kimi-k3
EmbeddingOpen weights

We present a sentence transformation model that maps sentences and paragraphs to a 768-dimensional dense vector space, suitable for semantic search tasks. The model is trained on 215 million question-answer pairs from various sources, including WikiAnswers, PAQ, Stack Exchange, MS MARCO, GOOAQ, Amazon QA, Yahoo Answers, Search QA, ELI5, and Natural Questions. Our model uses a contrastive learning objective.

Context

512

In EGP / 1M

0.28

Out EGP / 1M

Free

EmbeddingOpen weights

The Multilingual-E5-large model is a 24-layer text embedding model with an embedding size of 1024, trained on a mixture of multilingual datasets and supporting 100 languages.

Context

512

In EGP / 1M

0.56

Out EGP / 1M

Free

EmbeddingOpen weights

The Multilingual-E5 models, initialized from XLM-RoBERTa, support up to 512 tokens per input — any longer text will be silently truncated. To ensure optimal performance, always prefix inputs with “query:” or “passage:”, as the model was explicitly trained with this format.

Context

512

In EGP / 1M

0.56

Out EGP / 1M

Free

meta-models
ChatOpen weights

Context

131K

In EGP / 1M

19.60

Out EGP / 1M

84.01

muse-glimmer-30b
ChatOpen weights

NVIDIA Nemotron 3 Nano is an open small reasoning model optimized for fast, cost-efficient inference in agentic and production workloads. Built with a hybrid Mixture-of-Experts (MoE) and Mamba-Transformer architecture, it delivers strong multi-step reasoning, high token throughput, stable latency with predictable cost, and efficient deployment for agent-based systems. Designed for real-world AI systems where reasoning can generate significantly more tokens per prompt, Nemotron Nano reduces compute cost while maintaining strong reasoning quality.

Context

262K

In EGP / 1M

2.80

Out EGP / 1M

11.20

ChatOpen weights

Nemotron Content Safety 3.5 is a multimodal safety classifier developed by NVIDIA. A compact safety model that handles text, images, and custom policies. It outputs a safe/unsafe classification plus a reasoning trace, and can be used as an inference-time guardrail, as a judge for LLM safety testing and evaluation, or with the accompanying training dataset to post-train models for safer behavior.

Context

131K

In EGP / 1M

11.20

Out EGP / 1M

11.20

Meta's latest class of model (Llama 3.1) launched with a variety of sizes & flavors. This 70B instruct-tuned version is optimized for high quality dialogue usecases. It has demonstrated strong

Context

131K

In EGP / 1M

22.40

Out EGP / 1M

22.40

llama-3.1-70b-instruct

Meta's latest class of model (Llama 3.1) launched with a variety of sizes & flavors. This 8B instruct-tuned version is fast and efficient. It has demonstrated strong performance compared to

Context

131K

In EGP / 1M

2.80

Out EGP / 1M

4.48

llama-3.1-8b-instruct

The Meta Llama 3.3 multilingual large language model (LLM) is a pretrained and instruction tuned generative model in 70B (text in/text out). The Llama 3.3 instruction tuned text only model

Context

131K

In EGP / 1M

5.60

Out EGP / 1M

17.92

llama-3.3-70b-instruct

NVIDIA Nemotron 3 Ultra is an open frontier-reasoning and orchestration model from NVIDIA, with 55B active parameters out of 550B total (MoE). Built on a hybrid Transformer-Mamba mixture-of-experts architecture, it

Context

256K

In EGP / 1M

35.00

Out EGP / 1M

175.02

nemotron-3-ultra-550b-a55b
ChatOpen weights

NVIDIA Nemotron 3 Super is a hybrid Mixture-of-Experts (MoE) model engineered for highest compute efficiency and accuracy in multi-agent applications and specialized agentic systems. It is optimized to run many collaborating agents per application on a single GPU, delivering high accuracy for reasoning, tool use, and instruction following.

Context

262K

In EGP / 1M

4.76

Out EGP / 1M

22.40

ChatOpen weights

Nemotron 3 Ultra is built for, frontier reasoning, orchestration, coding agents, deep research, and complex enterprise workflows. It delivers up to 5x faster inference and up to 30% lower cost for agentic workloads while supporting up to 1M token context.

Context

262K

In EGP / 1M

28.00

Out EGP / 1M

123.21

ChatOpen weights

NVIDIA Nemotron 3.5 Lightning is NVIDIA's fastest open model for always-on agents and high-volume specialized tasks. It delivers a substantial leap in agentic capability over its predecessor Nemotron 3 Nano, with up to 4x higher throughput on a 1M-token context.

Context

262K

In EGP / 1M

4.48

Out EGP / 1M

11.20

o1

Live
OpenAI
Chat

The latest and strongest model family from OpenAI, o1 is designed to spend more time thinking before responding. The o1 model series is trained with large-scale reinforcement learning to reason

Context

200K

In EGP / 1M

840.10

Out EGP / 1M

3360.39

o3

Live
OpenAI
Chat

o3 is a well-rounded and powerful model across domains. It sets a new standard for math, science, coding, and visual reasoning tasks. It also excels at technical writing and instruction-following

Context

200K

In EGP / 1M

112.01

Out EGP / 1M

448.05

OpenAI
Chat

OpenAI o3-mini is a cost-efficient language model optimized for STEM reasoning tasks, particularly excelling in science, mathematics, and coding. This model supports the `reasoning_effort` parameter, which can be set to

Context

200K

In EGP / 1M

61.61

Out EGP / 1M

246.43

o3-pro

Live
OpenAI
Chat

The o-series of models are trained with reinforcement learning to think before they answer and perform complex reasoning. The o3-pro model uses more compute to think harder and provide consistently

Context

200K

In EGP / 1M

1120.13

Out EGP / 1M

4480.52

OpenAI
Chat

OpenAI o4-mini is a compact reasoning model in the o-series, optimized for fast, cost-efficient performance while retaining strong multimodal and agentic capabilities. It supports tool use and demonstrates competitive reasoning

Context

200K

In EGP / 1M

61.61

Out EGP / 1M

246.43

EmbeddingOpen weights

We present a sentence similarity model based on the Sentence Transformers architecture, which maps sentences to a 384-dimensional dense vector space. The model uses a pre-trained BERT encoder and applies mean pooling on top of the contextualized word embeddings to obtain sentence embeddings. We evaluate the model on the Sentence Embeddings Benchmark.

Context

512

In EGP / 1M

0.28

Out EGP / 1M

Free

phi-4

Live
Microsoft
ChatOpen weights

Phi-4 is a model built upon a blend of synthetic datasets, data from filtered public domain websites, and acquired academic books and Q&A datasets. The goal of this approach was to ensure that small capable models were trained with data focused on high quality and advanced reasoning.

Context

16K

In EGP / 1M

3.92

Out EGP / 1M

7.84

Qwen3 Coder Flash is Alibaba's fast and cost efficient version of their proprietary Qwen3 Coder Plus. It is a powerful coding agent model specializing in autonomous programming via tool calling

Context

1M

In EGP / 1M

10.92

Out EGP / 1M

54.61

qwen3-coder-flash
Chat

Qwen3-Coder-Next is an open-weight causal language model optimized for coding agents and local development workflows. It uses a sparse MoE design with 80B total parameters and only 3B activated per

Context

262K

In EGP / 1M

6.72

Out EGP / 1M

44.81

qwen3-coder-next

The Qwen3.5 native vision-language series Plus models are built on a hybrid architecture that integrates linear attention mechanisms with sparse mixture-of-experts models, achieving higher inference efficiency. In a variety of

Context

1M

In EGP / 1M

14.56

Out EGP / 1M

87.37

qwen3.5-plus-02-15

Qwen3.5 Plus (April 2026) is a large-scale multimodal language model from Alibaba. It accepts text, image, and video input and produces text output, with a 1M token context window. This

Context

1M

In EGP / 1M

16.80

Out EGP / 1M

100.81

qwen3.5-plus-20260420
Chat

The Qwen3.5 native vision-language Flash models are built on a hybrid architecture that integrates a linear attention mechanism with a sparse mixture-of-experts model, achieving higher inference efficiency. Compared to the

Context

1M

In EGP / 1M

3.64

Out EGP / 1M

14.56

qwen3.5-flash-02-23
Chat

Qwen3.6 Flash is a fast, efficient language model from Alibaba's Qwen 3.6 series. It supports text, image, and video input with a 1M token context window. Tiered pricing kicks in

Context

1M

In EGP / 1M

10.50

Out EGP / 1M

63.01

qwen3.6-flash

Qwen3.6-Max-Preview is a proprietary frontier model from Alibaba Cloud built on a sparse mixture-of-experts architecture with approximately 1 trillion total parameters. It is optimized for agentic coding, tool use, and

Context

262K

In EGP / 1M

57.52

Out EGP / 1M

345.11

qwen3.6-max-preview
Chat

Qwen3.7 Flash is a vision-language reasoning model from Alibaba. It is suited for multimodal agents, visual coding, search, and computer interaction, with strengths in object recognition, spatial understanding, and real-world

Context

1M

In EGP / 1M

1.68

Out EGP / 1M

7.28

qwen3.7-flash
ChatOpen weights

Qwen3.8 27B is an open-weight dense vision-language model from Qwen. It is suited for coding, professional workflows, research, multimodal interaction, and long-running agent tasks, with flexible thinking that can be

Context

262K

In EGP / 1M

22.40

Out EGP / 1M

168.02

qwen3.8-27b
Chat

Qwen3.8 Flash is a multimodal reasoning model from Alibaba. It is suited for coding assistance, agentic workflows, visual understanding, document and codebase analysis, desktop interaction, chart analysis, and long-video analysis.

Context

1M

In EGP / 1M

8.40

Out EGP / 1M

26.32

qwen3.8-flash
ChatProprietary

Qwen3.8 Max is the flagship model in Alibaba's Qwen3.8 series, the general-availability successor to the Qwen3.8 Max Preview. It is a multimodal reasoning model intended for complex reasoning, visual understanding,

Context

256K

In EGP / 1M

92.41

Out EGP / 1M

277.29

qwen3.8-max
ChatOpen weights

Qwen2.5-VL is proficient in recognizing common objects such as flowers, birds, fish, and insects. It is also highly capable of analyzing texts, charts, icons, graphics, and layouts within images.

Context

128K

In EGP / 1M

44.81

Out EGP / 1M

56.01

qwen2.5-vl-72b-instruct
Alibaba
ChatOpen weights

Qwen3-8B is a dense 8.2B parameter causal language model from the Qwen3 series, designed for both reasoning-heavy tasks and efficient dialogue. It supports seamless switching between "thinking" mode for math,

Context

131K

In EGP / 1M

6.55

Out EGP / 1M

25.48

qwen3-8b
ChatOpen weights

Qwen3-Coder-30B-A3B-Instruct is a 30.5B parameter Mixture-of-Experts (MoE) model with 128 experts (8 active per forward pass), designed for advanced code generation, repository-scale understanding, and agentic tool use. Built on the

Context

262K

In EGP / 1M

3.92

Out EGP / 1M

15.68

qwen3-coder-30b-a3b-instruct
ChatOpen weights

Qwen3-Next-80B-A3B-Thinking is a reasoning-first chat model in the Qwen3-Next line that outputs structured “thinking” traces by default. It’s designed for hard multi-step problems; math proofs, code synthesis/debugging, logic, and agentic

Context

262K

In EGP / 1M

8.40

Out EGP / 1M

67.21

qwen3-next-80b-a3b-thinking
Alibaba
ChatOpen weights

Qwen3 is the latest generation of large language models in Qwen series, offering a comprehensive suite of dense and mixture-of-experts (MoE) models. Built upon extensive training, Qwen3 delivers groundbreaking advancements in reasoning, instruction-following, agent capabilities, and multilingual support.

Context

41K

In EGP / 1M

6.72

Out EGP / 1M

13.44

ChatOpen weights

Qwen3-235B-A22B-Instruct-2507 is the updated version of the Qwen3-235B-A22B non-thinking mode, featuring Significant improvements in general capabilities, including instruction following, logical reasoning, text comprehension, mathematics, science, coding and tool usage.

Context

262K

In EGP / 1M

5.04

Out EGP / 1M

30.80

ChatOpen weights

Qwen3-235B-A22B-Thinking-2507 is the Qwen3's new model with scaling the thinking capability of Qwen3-235B-A22B, improving both the quality and depth of reasoning.

Context

131K

In EGP / 1M

12.88

Out EGP / 1M

128.81

Alibaba
ChatOpen weights

Qwen3 is the latest generation of large language models in Qwen series, offering a comprehensive suite of dense and mixture-of-experts (MoE) models. Built upon extensive training, Qwen3 delivers groundbreaking advancements in reasoning, instruction-following, agent capabilities, and multilingual support

Context

41K

In EGP / 1M

6.72

Out EGP / 1M

28.00

Alibaba
ChatOpen weights

Qwen3 is the latest generation of large language models in Qwen series, offering a comprehensive suite of dense and mixture-of-experts (MoE) models. Built upon extensive training, Qwen3 delivers groundbreaking advancements in reasoning, instruction-following, agent capabilities, and multilingual support

Context

41K

In EGP / 1M

4.48

Out EGP / 1M

15.68

EmbeddingOpen weights

The Qwen3 Embedding model series is the latest proprietary model of the Qwen family, specifically designed for text embedding and ranking tasks. Building upon the dense foundational models of the Qwen3 series, it provides a comprehensive range of text embeddings and reranking models in various sizes (0.6B, 4B, and 8B).

Context

33K

In EGP / 1M

0.56

Out EGP / 1M

Free

EmbeddingOpen weights

The Qwen3 Embedding model series is the latest proprietary model of the Qwen family, specifically designed for text embedding and ranking tasks. Building upon the dense foundational models of the Qwen3 series, it provides a comprehensive range of text embeddings and reranking models in various sizes (0.6B, 4B, and 8B).

Context

33K

In EGP / 1M

1.12

Out EGP / 1M

Free

EmbeddingOpen weights

The Qwen3 Embedding model series is the latest proprietary model of the Qwen family, specifically designed for text embedding and ranking tasks. Building upon the dense foundational models of the Qwen3 series, it provides a comprehensive range of text embeddings and reranking models in various sizes (0.6B, 4B, and 8B).

Context

33K

In EGP / 1M

0.56

Out EGP / 1M

Free

Alibaba
ChatProprietary

The latest flagship model in the Qwen family. State-of-the-art results across a comprehensive suite of benchmarks — including knowledge, reasoning, coding, instruction following, human preference alignment, agent tasks, and multilingual understanding.

Context

256K

In EGP / 1M

67.21

Out EGP / 1M

336.04

ChatProprietary

The latest flagship reasoning model in the Qwen3 family. Further enhanced by multiple innovations like adaptive tool-use and advanced test-time scaling techniques

Context

256K

In EGP / 1M

67.21

Out EGP / 1M

336.04

ChatOpen weights

Over the past few months, we have observed increasingly clear trends toward scaling both total parameters and context lengths in the pursuit of more powerful and agentic artificial intelligence (AI). We are excited to share our latest advancements in addressing these demands, centered on improving scaling efficiency through innovative model architecture. We call this next-generation foundation models Qwen3-Next.

Context

262K

In EGP / 1M

5.04

Out EGP / 1M

61.61

ChatOpen weights

Meet Qwen3-VL — the most powerful vision-language model in the Qwen series to date. This generation delivers comprehensive upgrades across the board: superior text understanding & generation, deeper visual perception & reasoning, extended context length, enhanced spatial and video dynamics comprehension, and stronger agent interaction capabilities.

Context

262K

In EGP / 1M

11.20

Out EGP / 1M

49.29

ChatOpen weights

Meet Qwen3-VL — the most powerful vision-language model in the Qwen series to date. This generation delivers comprehensive upgrades across the board: superior text understanding & generation, deeper visual perception & reasoning, extended context length, enhanced spatial and video dynamics comprehension, and stronger agent interaction capabilities.

Context

262K

In EGP / 1M

8.40

Out EGP / 1M

33.60

ChatOpen weights

Qwen3-VL-32B-Instruct is a large-scale multimodal vision-language model designed for high-precision understanding and reasoning across text, images, and video. With 32 billion parameters, it combines deep visual perception with advanced text

Context

131K

In EGP / 1M

5.82

Out EGP / 1M

23.30

qwen3-vl-32b-instruct
ChatOpen weights

Qwen3-VL-8B-Instruct is a multimodal vision-language model from the Qwen3-VL series, built for high-fidelity understanding and reasoning across text, images, and video. It features improved multimodal fusion with Interleaved-MRoPE for long-horizon

Context

131K

In EGP / 1M

6.55

Out EGP / 1M

25.48

qwen3-vl-8b-instruct
ChatOpen weights

Qwen3.5-122B-A10B is a large Mixture-of-Experts model from Alibaba's Qwen3.5 series with 122B total parameters and 10B activated per token. It features a 262K token context window (extensible to 1M with YaRN), thinking/reasoning mode, tool calling, and support for 201 languages. Excels at complex reasoning, coding, multimodal understanding, and agentic tasks with the efficiency of sparse activation.

Context

262K

In EGP / 1M

16.24

Out EGP / 1M

134.42

Alibaba
ChatOpen weights

Qwen3.5-27B is Alibaba's largest dense Qwen3.5 model, delivering near-frontier quality across reasoning, coding, and instruction following. It features a 262K token context window (extensible to 1M), thinking/reasoning mode, tool calling, multi-token prediction, and support for 201 languages. Best suited for production deployments and complex enterprise tasks requiring top-tier performance.

Context

262K

In EGP / 1M

14.56

Out EGP / 1M

145.62

Alibaba
ChatOpen weights

Qwen3.5-35B-A3B is an efficient Mixture-of-Experts model from Alibaba's Qwen3.5 series with 35B total parameters and only 3B activated per token. It features a 262K token context window (extensible to 1M with YaRN), thinking/reasoning mode, tool calling, and support for 201 languages. Delivers strong performance on reasoning, coding, and vision-language tasks at a fraction of the compute cost.

Context

262K

In EGP / 1M

7.84

Out EGP / 1M

56.01

ChatOpen weights

Qwen3.5-397B-A17B is Alibaba's most capable Qwen3.5 model, a Mixture-of-Experts architecture with 397B total parameters and 17B activated per token. It features a 262K token context window (extensible to 1M with YaRN), thinking/reasoning mode, tool calling with MCP integration, and support for 201 languages. Sets state-of-the-art results on reasoning, coding, math, and multimodal benchmarks.

Context

262K

In EGP / 1M

25.20

Out EGP / 1M

168.02

Alibaba
ChatOpen weights

Qwen3.5-9B is a high-performance model from Alibaba's Qwen3.5 series with a hybrid Gated Delta Networks and sparse MoE architecture. It features a 262K token context window, thinking/reasoning mode, tool calling, multi-token prediction, and support for 201 languages. Excels at reasoning, coding, instruction following, and long-context tasks.

Context

262K

In EGP / 1M

5.60

Out EGP / 1M

8.40

Alibaba
Chat

Qwen 3.6 Plus builds on a hybrid architecture that combines efficient linear attention with sparse mixture-of-experts routing, enabling strong scalability and high-performance inference. Compared to the 3.5 series, it delivers

Context

1M

In EGP / 1M

28.00

Out EGP / 1M

168.02

qwen3.6-plus
Alibaba
ChatOpen weights

Context

262K

In EGP / 1M

17.92

Out EGP / 1M

179.22

Alibaba
ChatOpen weights

Qwen3.6-35B-A3B is Alibaba's latest flagship Mixture-of-Experts model, with 35B total parameters and only 3B activated per token (256 experts, 8 routed + 1 shared). Built on direct feedback from the community, Qwen3.6 prioritizes stability and real-world utility, offering developers a more intuitive, responsive, and genuinely productive coding experience.

Context

262K

In EGP / 1M

5.60

Out EGP / 1M

53.21

Alibaba
ChatProprietary

Qwen3.7-Max is the flagship model in Alibaba's Qwen3.7 series. It supports text input and output and is designed for agent-centric workloads, with particular strengths in coding, office and productivity tasks,

Context

256K

In EGP / 1M

140.02

Out EGP / 1M

420.05

qwen3.7-max
Alibaba
Chat

Qwen3.7-Plus is a cost-effective model in Alibaba's Qwen3.7 series. It supports text and image input with text output, building on the series' text capabilities with a comprehensive upgrade to its

Context

1M

In EGP / 1M

17.92

Out EGP / 1M

71.69

qwen3.7-plus
ChatOpen weights

Qwen3.8 2.4T A95B is an open-weight sparse mixture-of-experts model from Qwen and the open-weight variant of Qwen3.8 Max, with 95 billion active parameters out of 2.4 trillion total. It is suited for coding, research, complex reasoning, and agentic workflows.

Context

262K

In EGP / 1M

112.01

Out EGP / 1M

336.04

ByteDance
ChatProprietary

Optimized specifically for multimodal agent scenarios. It features enhanced agent capabilities, upgraded multimodal comprehension, and more flexible context management.

Context

256K

In EGP / 1M

14.00

Out EGP / 1M

112.01

ByteDance
ChatProprietary

A coding model optimized for real-world development environments, with reliable tool use in common IDEs such as Claude Code. It delivers strong front-end performance and supports Skills.

Context

256K

In EGP / 1M

28.00

Out EGP / 1M

168.02

ByteDance
ChatProprietary

Built for low-latency, high-concurrency, cost-sensitive use cases, with flexible deployment, four-tier thinking, and multimodal

Context

256K

In EGP / 1M

5.60

Out EGP / 1M

22.40

ByteDance
ChatProprietary

Built for the Agent era, it delivers stable performance in complex reasoning and long-horizon tasks, including multi-step planning, visual-text reasoning, video understanding, and advanced analysis.

Context

256K

In EGP / 1M

28.00

Out EGP / 1M

168.02

stepfun-ai
ChatOpen weights

Step 3.7 Flash is an open-source multimodal reasoning model by StepFun with 198B total parameters (11B active) using Mixture of Experts. It accepts text and image inputs and features a 256K context window, selectable reasoning effort, tool calling, and agentic capabilities for coding and search workflows, scoring 80.9% on GPQA Diamond and 56.3% on SWE-bench Pro.

Context

262K

In EGP / 1M

11.20

Out EGP / 1M

64.41

Embedding

Context

In EGP / 1M

7.28

Out EGP / 1M

Free

Embedding

Context

In EGP / 1M

1.12

Out EGP / 1M

Free

Embedding

Context

In EGP / 1M

5.60

Out EGP / 1M

Free

shibing624
EmbeddingOpen weights

A sentence similarity model that can be used for various NLP tasks such as text classification, sentiment analysis, named entity recognition, question answering, and more. It utilizes the CoSENT architecture, which consists of a transformer encoder and a pooling module, to encode input texts into vectors that capture their semantic meaning. The model was trained on the nli_zh dataset and achieved high performance on various benchmark datasets.

Context

512

In EGP / 1M

0.28

Out EGP / 1M

Free

thinkingmachines
ChatOpen weights

Inkling Small is an open-weight multimodal mixture-of-experts model from Thinking Machines Lab, with 12B active parameters out of 276B total. It is positioned as the smaller, more efficient member of

Context

524K

In EGP / 1M

25.20

Out EGP / 1M

67.21

inkling-small
ChatOpen weights

Voxtral Small is an enhancement of Mistral Small 3, incorporating state-of-the-art audio input capabilities while retaining best-in-class text performance. It excels at speech transcription, translation and audio understanding.

Context

33K

In EGP / 1M

5.60

Out EGP / 1M

16.80

z-ai
ChatOpen weights

GLM-5.3 is a large-scale reasoning model from Z.ai, built for complex software engineering and long-horizon agent tasks. It supports text input and output with a 1M-token context window, and improves

Context

1M

In EGP / 1M

67.21

Out EGP / 1M

224.03

glm-5.3