Model Library · US

402 models

Available models with live pricing, context windows, and status.

More filters
Max input priceAnyEGP / 1M
Max output priceAnyEGP / 1M

Context window

Providers

402 models matching
aion-labs
Chat

Aion 3.5 is a multi-model roleplaying and storytelling system from AionLabs, built on the GLM family of models. It uses a collaborative generation process in which multiple specialized models each

Context

262K

In EGP / 1M

171.90

Out EGP / 1M

343.81

aion-3.5
Chat

Aion 3.5 Mini is a multi-model roleplaying and storytelling system from AionLabs, built on the GLM family of models. It is the smaller, lower-cost sibling of Aion 3.5 and uses

Context

262K

In EGP / 1M

40.11

Out EGP / 1M

80.22

aion-3.5-mini
aion-labs
Chat

Aion-2.0 is a variant of DeepSeek V3.2 optimized for immersive roleplaying and storytelling. It is particularly strong at introducing tension, crises, and conflict into stories, making narratives feel more engaging

Context

131K

In EGP / 1M

45.84

Out EGP / 1M

91.68

aion-2.0
aion-labs
Chat

Aion-3.0 is a multi-model roleplaying and storytelling system from AionLabs, built on the GLM family of models. It uses a collaborative generation process in which multiple specialized models each contribute

Context

131K

In EGP / 1M

171.90

Out EGP / 1M

343.81

aion-3.0
Chat

Aion-3.0 Mini is a multi-model roleplaying and storytelling system from AionLabs, built on the DeepSeek family of models. It uses a collaborative generation process in which multiple specialized models each

Context

131K

In EGP / 1M

40.11

Out EGP / 1M

80.22

aion-3.0-mini

Aion-RP-Llama-3.1-8B ranks the highest in the character evaluation portion of the RPBench-Auto benchmark, a roleplaying-specific variant of Arena-Hard-Auto, where LLMs evaluate each other’s responses. It is a fine-tuned base model

Context

33K

In EGP / 1M

45.84

Out EGP / 1M

91.68

aion-rp-llama-3.1-8b
EmbeddingOpen weights

We present a sentence transformation model that generates semantically similar sentences. Our model is based on the Sentence-Transformers architecture and was trained on a large dataset of sentence pairs. We evaluate the effectiveness of our model by measuring its ability to generate similar sentences that are close to the original sentence in meaning.

Context

512

In EGP / 1M

0.29

Out EGP / 1M

—

EmbeddingOpen weights

We present a sentence transformation model that achieves state-of-the-art results on various NLP tasks without requiring task-specific architectures or fine-tuning. Our approach leverages contrastive learning and utilizes a variety of datasets to learn robust sentence representations. We evaluate our model on several benchmarks and demonstrate its effectiveness in various applications such as text classification, sentiment analysis, named entity recognition, and question answering.

Context

512

In EGP / 1M

0.29

Out EGP / 1M

—

EmbeddingOpen weights

A sentence transformation model that has been trained on a wide range of datasets, including but not limited to S2ORC, WikiAnwers, PAQ, Stack Exchange, and Yahoo! Answers. Our model can be used for various NLP tasks such as clustering, sentiment analysis, and question answering.

Context

512

In EGP / 1M

0.29

Out EGP / 1M

—

Chat

Nova 2 Lite is a fast, cost-effective reasoning model for everyday workloads that can process text, images, and videos to generate text. Nova 2 Lite demonstrates standout capabilities in processing

Context

1M

In EGP / 1M

17.19

Out EGP / 1M

143.25

nova-2-lite-v1
Chat

Amazon Nova Lite 1.0 is a very low-cost multimodal model from Amazon that focused on fast processing of image, video, and text inputs to generate text output. Amazon Nova Lite

Context

300K

In EGP / 1M

3.44

Out EGP / 1M

13.75

nova-lite-v1
Chat

Amazon Nova Micro 1.0 is a text-only model that delivers the lowest latency responses in the Amazon Nova family of models at a very low cost. With a context length

Context

128K

In EGP / 1M

2.01

Out EGP / 1M

8.02

nova-micro-v1
Chat

Amazon Nova Pro 1.0 is a capable multimodal model from Amazon focused on providing a combination of accuracy, speed, and cost for a wide range of tasks. As of December

Context

300K

In EGP / 1M

45.84

Out EGP / 1M

183.36

nova-pro-v1
ChatProprietary

Claude Fable 5 is a Mythos-class model from Anthropic, built for autonomous knowledge work and coding. It supports text, image, and file inputs with text output, with reasoning support and

Context

1M

In EGP / 1M

573.02

Out EGP / 1M

2865.08

claude-fable-5

Claude Fable 5.1 improves on Claude Fable 5 across the board, with the biggest gains in agentic coding, long-running agentic workflows, and knowledge work: long code refactors, front-end and visual

Context

1M

In EGP / 1M

573.02

Out EGP / 1M

2865.08

claude-fable-5.1

This model always redirects to the latest model in the Claude Fable family.

Context

1M

In EGP / 1M

573.02

Out EGP / 1M

2865.08

claude-fable-latest

This model always redirects to the latest model in the Claude Haiku family.

Context

200K

In EGP / 1M

57.30

Out EGP / 1M

286.51

claude-haiku-latest

Claude Opus 4.1 is an updated version of Anthropic’s flagship model, offering improved performance in coding, reasoning, and agentic tasks. It achieves 74.5% on SWE-bench Verified and shows notable gains

Context

200K

In EGP / 1M

859.52

Out EGP / 1M

4297.61

claude-opus-4.1

Claude Opus 4.5 is Anthropic’s frontier reasoning model optimized for complex software engineering, agentic workflows, and long-horizon computer use. It offers strong multimodal capabilities, competitive performance across real-world coding and

Context

200K

In EGP / 1M

286.51

Out EGP / 1M

1432.54

claude-opus-4.5

Opus 4.6 is Anthropic’s strongest model for coding and long-running professional tasks. It is built for agents that operate across entire workflows rather than single prompts, making it especially effective

Context

1M

In EGP / 1M

286.51

Out EGP / 1M

1432.54

claude-opus-4.6
ChatProprietary

Opus 4.7 is the next generation of Anthropic's Opus family, built for long-running, asynchronous agents. Building on the coding and agentic strengths of Opus 4.6, it delivers stronger performance on

Context

1M

In EGP / 1M

286.51

Out EGP / 1M

1432.54

claude-opus-4.7
ChatProprietary

Claude Opus 4.8 is Anthropic's most capable generally available model in the Opus family. It supports text, image, and file inputs with text output, with reasoning support and a 1M-token

Context

1M

In EGP / 1M

286.51

Out EGP / 1M

1432.54

claude-opus-4.8
ChatProprietary

Claude Opus 5 is Anthropic’s flagship model for demanding reasoning, coding, and long-horizon agentic work. It is particularly strong at end-to-end software tasks, code review and bug finding, visual analysis

Context

1M

In EGP / 1M

286.51

Out EGP / 1M

1432.54

claude-opus-5
ChatProprietary

Claude Opus 5.5 is Anthropic's flagship model for demanding reasoning, coding, and long-horizon agentic work, succeeding Claude Opus 5. It is particularly strong at multi-step changes in large codebases, code

Context

1M

In EGP / 1M

229.21

Out EGP / 1M

1146.03

claude-opus-5.5

This model always redirects to the latest model in the Claude Opus family.

Context

1M

In EGP / 1M

229.21

Out EGP / 1M

1146.03

claude-opus-latest

Claude Sonnet 4 significantly enhances the capabilities of its predecessor, Sonnet 3.7, excelling in both coding and reasoning tasks with improved precision and controllability. Achieving state-of-the-art performance on SWE-bench (72.7%),

Context

200K

In EGP / 1M

171.90

Out EGP / 1M

859.52

claude-sonnet-4

Claude Sonnet 4.5 is Anthropic’s most advanced Sonnet model to date, optimized for real-world agents and coding workflows. It delivers state-of-the-art performance on coding benchmarks such as SWE-bench Verified, with

Context

1M

In EGP / 1M

171.90

Out EGP / 1M

859.52

claude-sonnet-4.5
ChatProprietary

Sonnet 4.6 is Anthropic's most capable Sonnet-class model yet, with frontier performance across coding, agents, and professional work. It excels at iterative development, complex codebase navigation, end-to-end project management with

Context

1M

In EGP / 1M

171.90

Out EGP / 1M

859.52

claude-sonnet-4.6
ChatProprietary

Sonnet 5 is Anthropic's most capable Sonnet-class model, with frontier performance across coding, agents, and professional work. It supports adaptive thinking with selectable reasoning effort levels (low, medium, high, max,

Context

1M

In EGP / 1M

171.90

Out EGP / 1M

859.52

claude-sonnet-5

Claude Sonnet 5.5 is Anthropic's Sonnet-class model for well-scoped everyday work, succeeding Claude Sonnet 5 as a direct upgrade. It is especially strong at building features, fixing bugs, and producing

Context

1M

In EGP / 1M

114.60

Out EGP / 1M

573.02

claude-sonnet-5.5

This model always redirects to the latest model in the Claude Sonnet family.

Context

1M

In EGP / 1M

114.60

Out EGP / 1M

573.02

claude-sonnet-latest

Trinity Large Thinking is a powerful open source reasoning model from the team at Arcee AI. It shows strong performance in PinchBench, agentic workloads, and reasoning tasks. Launch video: https://youtu.be/Gc82AXLa0Rg?si=4RLn6WBz33qT--B7

Context

262K

In EGP / 1M

14.33

Out EGP / 1M

45.84

trinity-large-thinking

ERNIE-4.5-VL-424B-A47B is a multimodal Mixture-of-Experts (MoE) model from Baidu’s ERNIE 4.5 series, featuring 424B total parameters with 47B active per token. It is trained jointly on text and image data

Context

123K

In EGP / 1M

24.07

Out EGP / 1M

71.63

ernie-4.5-vl-424b-a47b
EmbeddingOpen weights

BGE embedding is a general Embedding Model. It is pre-trained using retromae and trained on large-scale pair data using contrastive learning. Note that the goal of pre-training is to reconstruct the text, and the pre-trained model cannot be used for similarity calculation directly, it needs to be fine-tuned

Context

512

In EGP / 1M

0.29

Out EGP / 1M

—

BAAI
EmbeddingOpen weights

A LLM-based embedding model with in-context learning capabilities that achieves SOTA performance on BEIR and AIR-Bench. It leverages few-shot examples to enhance task performance.

Context

8K

In EGP / 1M

0.57

Out EGP / 1M

—

EmbeddingOpen weights

BGE embedding is a general Embedding Model. It is pre-trained using retromae and trained on large-scale pair data using contrastive learning. Note that the goal of pre-training is to reconstruct the text, and the pre-trained model cannot be used for similarity calculation directly, it needs to be fine-tuned

Context

512

In EGP / 1M

0.57

Out EGP / 1M

—

bge-m3

Live
BAAI
EmbeddingOpen weights

BGE-M3 is a versatile text embedding model that supports multi-functionality, multi-linguality, and multi-granularity, allowing it to perform dense retrieval, multi-vector retrieval, and sparse retrieval in over 100 languages and with input sizes up to 8192 tokens. The model can be used in a retrieval pipeline with hybrid retrieval and re-ranking to achieve higher accuracy and stronger generalization capabilities. BGE-M3 has shown state-of-the-art performance on several benchmarks, including MKQA, MLDR, and NarritiveQA, and can be used as a drop-in replacement for other embedding models like DPR and BGE-v1.5.

Context

8K

In EGP / 1M

0.57

Out EGP / 1M

—

BAAI
EmbeddingOpen weights

BGE-M3 is a multilingual text embedding model developed by BAAI, distinguished by its Multi-Linguality (supporting 100+ languages), Multi-Functionality (unified dense, multi-vector, and sparse retrieval), and Multi-Granularity (handling inputs from short queries to long documents). It achieves state-of-the-art retrieval performance across diverse benchmarks while maintaining a single model for multiple retrieval modes. Inputs to this endpoint are truncated to 512 tokens.

Context

512

In EGP / 1M

0.57

Out EGP / 1M

—

EmbeddingOpen weights

BGE-M3 is a multilingual text embedding model developed by BAAI, distinguished by its Multi-Linguality (supporting 100+ languages), Multi-Functionality (unified dense, multi-vector, and sparse retrieval), and Multi-Granularity (handling inputs from short queries to long documents). It achieves state-of-the-art retrieval performance across diverse benchmarks while maintaining a single model for multiple retrieval modes. This endpoint serves the model's full 8192-token context.

Context

8K

In EGP / 1M

0.57

Out EGP / 1M

—

Bria
imageProprietary

Bria 3.2 is the next-generation commercial-ready text-to-image model. With just 4 billion parameters, it provides exceptional aesthetics and text rendering, evaluated to be on par to leading open-source models, and outperforming other licensed models.

Context

—

Price

2.29 EGP per image

imageProprietary

Bria 3.2 is the next-generation commercial-ready text-to-image model. With just 4 billion parameters, it provides exceptional aesthetics and text rendering, evaluated to be on par to leading open-source models, and outperforming other licensed models.

Context

—

Price

2.29 EGP per image

bytedance-seed
Chat

Seed 1.6 is a general-purpose model released by the ByteDance Seed team. It incorporates multimodal capabilities and adaptive deep thinking with a 256K context window.

Context

262K

In EGP / 1M

14.33

Out EGP / 1M

114.60

seed-1.6
Chat

Seed 1.6 Flash is an ultra-fast multimodal deep thinking model by ByteDance Seed, supporting both text and visual understanding. It features a 256k context window and can generate outputs of

Context

262K

In EGP / 1M

4.30

Out EGP / 1M

17.19

seed-1.6-flash
Chat

Seed 2.1 Turbo is a multimodal model from ByteDance Seed for coding and long-horizon agent workflows. It is suited for end-to-end software delivery, multi-step task execution, and understanding visual and

Context

262K

In EGP / 1M

28.65

Out EGP / 1M

143.25

seed-2-1-turbo
ChatProprietary

Seed 2.0 Code is a model from ByteDance Seed optimized for agentic coding. It is suited for frontend development, multilingual programming tasks, and coding-agent workflows in tools such as Claude

Context

256K

In EGP / 1M

28.65

Out EGP / 1M

171.90

seed-2.0-code
Chat

Seed-2.0-Lite is a versatile, cost‑efficient enterprise workhorse that delivers strong multimodal and agent capabilities while offering noticeably lower latency, making it a practical default choice for most production workloads across

Context

262K

In EGP / 1M

14.33

Out EGP / 1M

114.60

seed-2.0-lite
ChatProprietary

Seed-2.0-mini targets latency-sensitive, high-concurrency, and cost-sensitive scenarios, emphasizing fast response and flexible inference deployment. It delivers performance comparable to ByteDance-Seed-1.6, supports 256k context, four reasoning effort modes (minimal/low/medium/high), multimodal understanding,

Context

256K

In EGP / 1M

5.73

Out EGP / 1M

22.92

seed-2.0-mini
Chat

UI-TARS-1.5 is a multimodal vision-language agent optimized for GUI-based environments, including desktop interfaces, web browsers, mobile systems, and games. Built by ByteDance, it builds upon the UI-TARS framework with reinforcement

Context

128K

In EGP / 1M

5.73

Out EGP / 1M

11.46

ui-tars-1.5-7b
Anthropic
ChatProprietary

The next generation of Anthropic's fastest and most cost-effective model, optimal for use cases where speed and affordability matter.

Context

200K

In EGP / 1M

57.30

Out EGP / 1M

286.51

SBERT
EmbeddingOpen weights

The CLIP model maps text and images to a shared vector space, enabling various applications such as image search, zero-shot image classification, and image clustering. The model can be used easily after installation, and its performance is demonstrated through zero-shot ImageNet validation set accuracy scores. Multilingual versions of the model are also available for 50+ languages.

Context

77

In EGP / 1M

0.29

Out EGP / 1M

—

EmbeddingOpen weights

This model is a multilingual version of the OpenAI CLIP-ViT-B32 model, which maps text and images to a common dense vector space. It includes a text embedding model that works for 50+ languages and an image encoder from CLIP. The model was trained using Multilingual Knowledge Distillation, where a multilingual DistilBERT model was trained as a student model to align the vector space of the original CLIP image encoder across many languages.

Context

512

In EGP / 1M

0.29

Out EGP / 1M

—

cohere
Chat

Command A is an open-weights 111B parameter model with a 256k context window focused on delivering great performance across agentic, multilingual, and coding use cases. Compared to other leading proprietary

Context

256K

In EGP / 1M

143.25

Out EGP / 1M

573.02

command-a
cohere
Chat

Command A+ is Cohere's flagship model for enterprise agentic workflows. It accepts text and image inputs with a 192K context window, supports native tool calling with strict tool schemas, structured

Context

192K

In EGP / 1M

17.19

Out EGP / 1M

85.95

command-a-plus

command-r-08-2024 is an update of the Command R with improved performance for multilingual retrieval-augmented generation (RAG) and tool use. More broadly, it is better at math, code and reasoning and

Context

128K

In EGP / 1M

8.60

Out EGP / 1M

34.38

command-r-08-2024

command-r-plus-08-2024 is an update of the Command R+ with roughly 50% higher throughput and 25% lower latencies as compared to the previous Command R+ version, while keeping the hardware footprint

Context

128K

In EGP / 1M

143.25

Out EGP / 1M

573.02

command-r-plus-08-2024

Command R7B (12-2024) is a small, fast update of the Command R+ model, delivered in December 2024. It excels at RAG, tool use, agents, and similar tasks requiring complex reasoning

Context

128K

In EGP / 1M

2.15

Out EGP / 1M

8.60

command-r7b-12-2024
DeepSeek
ChatOpen weights

The DeepSeek R1 model has undergone a minor version upgrade, with the current version being DeepSeek-R1-0528.

Context

164K

In EGP / 1M

28.65

Out EGP / 1M

123.20

DeepSeek
ChatOpen weights

DeepSeek-V3, a strong Mixture-of-Experts (MoE) language model with 671B total parameters with 37B activated for each token. To achieve efficient inference and cost-effective training, DeepSeek-V3 adopts Multi-head Latent Attention (MLA) and DeepSeekMoE architectures, which were thoroughly validated in DeepSeek-V2.

Context

164K

In EGP / 1M

18.34

Out EGP / 1M

51.00

DeepSeek
ChatOpen weights

DeepSeek-V3.1 is post-trained on the top of DeepSeek-V3.1-Base, which is built upon the original V3 base checkpoint through a two-phase long context extension approach, following the methodology outlined in the original DeepSeek-V3 report. We have expanded our dataset by collecting additional long documents and substantially extending both training phases. The 32K extension phase has been increased 10-fold to 630B tokens, while the 128K extension phase has been extended by 3.3x to 209B tokens. Additionally, DeepSeek-V3.1 is trained using the UE8M0 FP8 scale data format to ensure compatibility with microscaling data formats.

Context

164K

In EGP / 1M

14.33

Out EGP / 1M

54.44

This model always redirects to the latest model in the DeepSeek Flash family.

Context

1M

In EGP / 1M

1.07

Out EGP / 1M

22.69

deepseek-flash-latest

This model always redirects to the latest model in the DeepSeek Pro family.

Context

1M

In EGP / 1M

7.05

Out EGP / 1M

200.56

deepseek-pro-latest
Chat

DeepSeek-V3 is the latest model from the DeepSeek team, building upon the instruction following and coding abilities of the previous versions. Pre-trained on nearly 15 trillion tokens, the reported evaluations

Context

128K

In EGP / 1M

14.75

Out EGP / 1M

58.95

deepseek-chat

DeepSeek V3, a 685B-parameter, mixture-of-experts model, is the latest iteration of the flagship chat model family from the DeepSeek team. It succeeds the DeepSeek V3 model and performs really well

Context

164K

In EGP / 1M

16.62

Out EGP / 1M

65.32

deepseek-chat-v3-0324
Chat

DeepSeek-V3.1 is a large hybrid reasoning model (671B parameters, 37B active) that supports both thinking and non-thinking modes via prompt templates. It extends the DeepSeek-V3 base with a two-phase long-context

Context

164K

In EGP / 1M

14.33

Out EGP / 1M

54.44

deepseek-chat-v3.1

DeepSeek-V3.1 Terminus is an update to DeepSeek V3.1 that maintains the model's original capabilities while addressing issues reported by users, including language consistency and agent capabilities, further optimizing the model's

Context

131K

In EGP / 1M

17.19

Out EGP / 1M

57.30

deepseek-v3.1-terminus
ChatOpen weights

DeepSeek-V3.2 is a large language model designed to harmonize high computational efficiency with strong reasoning and agentic tool-use performance. It introduces DeepSeek Sparse Attention (DSA), a fine-grained sparse attention mechanism

Context

164K

In EGP / 1M

14.90

Out EGP / 1M

21.77

deepseek-v3.2

DeepSeek-V3.2-Exp is an experimental large language model released by DeepSeek as an intermediate step between V3.1 and future architectures. It introduces DeepSeek Sparse Attention (DSA), a fine-grained sparse attention mechanism

Context

164K

In EGP / 1M

15.47

Out EGP / 1M

23.49

deepseek-v3.2-exp
ChatOpen weights

DeepSeek V4 Flash is an efficiency-optimized Mixture-of-Experts model from DeepSeek with 284B total parameters and 13B activated parameters, supporting a 1M-token context window. It is designed for fast inference and

Context

1M

In EGP / 1M

5.16

Out EGP / 1M

10.31

deepseek-v4-flash
ChatOpen weights

DeepSeek V4 Flash 0731 is a sparse mixture-of-experts model from DeepSeek, with 13B active parameters out of 284B total. This re-post-trained revision is suited for coding, reasoning, and agent workflows

Context

1M

In EGP / 1M

3.44

Out EGP / 1M

10.31

deepseek-v4-flash-0731

This model always redirects to the latest model in the DeepSeek V4 Flash family.

Context

1M

In EGP / 1M

0.45

Out EGP / 1M

7.49

deepseek-v4-flash-latest
ChatOpen weights

DeepSeek V4 Flash Vision Exp is an experimental vision-enabled version of DeepSeek V4 Flash 0731 from DeepSeek, adding image understanding while matching the base model on text capabilities including agents,

Context

1M

In EGP / 1M

25.21

Out EGP / 1M

75.64

deepseek-v4-flash-vision-exp
ChatOpen weights

DeepSeek V4 Pro is a large-scale Mixture-of-Experts model from DeepSeek with 1.6T total parameters and 49B activated parameters, supporting a 1M-token context window. It is designed for advanced reasoning, coding,

Context

1M

In EGP / 1M

74.49

Out EGP / 1M

148.98

deepseek-v4-pro
ChatOpen weights

DeepSeek V4 Pro 0813 is a large-scale mixture-of-experts model from DeepSeek. This is the GA release of DeepSeek V4 Pro.

Context

1M

In EGP / 1M

74.49

Out EGP / 1M

148.98

deepseek-v4-pro-0813
ChatOpen weights

DeepSeek V4.1 Flash is a sparse mixture-of-experts model from DeepSeek, and the first built on the company's Causal Encoder-Decoder (CED) architecture. It activates 8B parameters on input and 16B on

Context

1M

In EGP / 1M

11.46

Out EGP / 1M

34.38

deepseek-v4.1-flash
DeepSeek
Chat

DeepSeek R1 is here: Performance on par with OpenAI o1, but open-sourced and with fully open reasoning tokens. It's 671B parameters in size, with 37B active in an inference pass

Context

64K

In EGP / 1M

40.11

Out EGP / 1M

143.25

deepseek-r1
intfloat
EmbeddingOpen weights

Text Embeddings by Weakly-Supervised Contrastive Pre-training. Model has 24 layers and 1024 out dim.

Context

512

In EGP / 1M

0.29

Out EGP / 1M

—

intfloat
EmbeddingOpen weights

Text Embeddings by Weakly-Supervised Contrastive Pre-training. Model has 24 layers and 1024 out dim.

Context

512

In EGP / 1M

0.57

Out EGP / 1M

—

EmbeddingOpen weights

EmbeddingGemma is a 300M parameter multilingual open embedding model from Google DeepMind, designed for efficient deployment even on low-resource devices, producing high-quality text vector representations for tasks such as search, classification, clustering, and semantic similarity.

Context

2K

In EGP / 1M

0.11

Out EGP / 1M

—

fibo

Live
Bria
imageProprietary

FIBO is an open-source, JSON-native text-to-image model trained on detailed structured descriptions (over 1,000+ words per image), providing fine-grained control over light, composition, and camera parameters.

Context

—

Price

2.29 EGP per image

Bria
imageProprietary

FIBO 1.5 is Bria's JSON-native text-to-image model, distilled to a few sampling steps and post-trained for sharper realism and texture, with structured-prompt control over lighting, composition and camera.

Context

—

Price

2.29 EGP per image

black-forest-labs
imageOpen weights

FLUX.1-dev is a state-of-the-art 12 billion parameter rectified flow transformer developed by Black Forest Labs. This model excels in text-to-image generation, providing highly accurate and detailed outputs. It is particularly well-regarded for its ability to follow complex prompts and generate anatomically accurate images, especially with challenging details like hands and faces.

Context

—

Price

0.5157 EGP per 1024p image

black-forest-labs
imageOpen weights

FLUX.1 [schnell] is a 12 billion parameter rectified flow transformer capable of generating images from text descriptions. This model offers cutting-edge output quality and competitive prompt following, matching the performance of closed source alternatives. Trained using latent adversarial diffusion distillation, FLUX.1 [schnell] can generate high-quality images in only 1 to 4 steps.

Context

—

Price

0.0287 EGP per 1024p image

black-forest-labs
imageProprietary

Black Forest Labs' latest state-of-the art proprietary model sporting top of the line prompt following, visual quality, details and output diversity.

Context

—

Price

2.29 EGP per image

black-forest-labs
imageOpen weights

Brand-new Flux2 Dev introduces a faster, more modular architecture for next-generation image generation pipelines. It delivers improved performance, cleaner control APIs, and a significantly more flexible development workflow for custom inference setups.

Context

—

Price

1.02 EGP per 1024p image

black-forest-labs
imageOpen weights

The fastest model of the Flux 2 family. Frontier visual intelligence — state-of-the-art image generation and editing from Black Forest Labs

Context

—

Price

0.8022 EGP per 1024p image

black-forest-labs
imageOpen weights

The best quality-to-latency ratio, production apps model of the Flux 2 family. Frontier visual intelligence — state-of-the-art image generation and editing from Black Forest Labs

Context

—

Price

0.8595 EGP per 1024p image

black-forest-labs
imageProprietary

The new top-tier image model from Black Forest Labs, significantly pushing image quality and editing consistency

Context

—

Price

≈ 4.01 EGP per image

black-forest-labs
imageProprietary

Multi-reference visual intelligence with unprecedented detail, color precision, and spatial reasoning. The most advanced image generation and editing model. Generate photorealistic images with precise control.

Context

—

Price

≈ 0.8595 EGP per image

ChatProprietary

Gemini 2.5 Flash is Google's latest thinking model, designed to tackle increasingly complex problems. It's capable of reasoning through their thoughts before responding, resulting in enhanced performance and improved accuracy. Gemini 2.5 Flash: best for balancing reasoning and speed.

Context

1M

In EGP / 1M

17.19

Out EGP / 1M

143.25

Google
ChatProprietary

Gemini 2.5 Pro is Google's the most advanced thinking model, designed to tackle increasingly complex problems. Gemini 2.5 Pro leads common benchmarks by meaningful margins and showcases strong reasoning and code capabilities. Gemini 2.5 models are thinking models, capable of reasoning through their thoughts before responding, resulting in enhanced performance and improved accuracy. The Gemini 2.5 Pro model is now available on DeepInfra.

Context

1M

In EGP / 1M

71.63

Out EGP / 1M

573.02

Google
ChatProprietary

Bring any idea to life with state-of-the-art reasoning to help you learn, build, and plan anything. Best for complex tasks and bringing creative concepts to life.

Context

1M

In EGP / 1M

114.60

Out EGP / 1M

687.62

Google
ChatOpen weights

Gemma 3 introduces multimodality, supporting vision-language input and text outputs. It handles context windows up to 128k tokens, understands over 140 languages, and offers improved math, reasoning, and chat capabilities, including structured outputs and function calling. Gemma 3-12B is Google's latest open source model, successor to Gemma 2

Context

131K

In EGP / 1M

2.87

Out EGP / 1M

8.60

Google
ChatOpen weights

Gemma 3 introduces multimodality, supporting vision-language input and text outputs. It handles context windows up to 128k tokens, understands over 140 languages, and offers improved math, reasoning, and chat capabilities, including structured outputs and function calling. Gemma 3 27B is Google's latest open source model, successor to Gemma 2

Context

131K

In EGP / 1M

4.58

Out EGP / 1M

9.17

ChatOpen weights

Gemma is a family of open models built by Google DeepMind. Gemma 4 models are multimodal, handling text and image input and generating text output.

Context

262K

In EGP / 1M

5.16

Out EGP / 1M

19.48

zai-org
ChatOpen weights

Compared with GLM-4.5, GLM-4.6 brings several key improvements: Longer context window: The context window has been expanded from 128K to 200K tokens, enabling the model to handle more complex agentic tasks. Superior coding performance: The model achieves higher scores on code benchmarks and demonstrates better real-world performance in applications such as Claude Code、Cline、Roo Code and Kilo Code, including improvements in generating visually polished front-end pages. Advanced reasoning: GLM-4.6 shows a clear improvement in reasoning performance and supports tool use during inference, leading to stronger overall capability. More capable agents: GLM-4.6 exhibits stronger performance in tool using and search-based agents, and integrates more effectively within agent frameworks. Refined writing: Better aligns with human preferences in style and readability, and performs more naturally in role-playing scenarios.

Context

203K

In EGP / 1M

28.65

Out EGP / 1M

114.60

Gemini 2.5 Flash-Lite is a lightweight reasoning model in the Gemini 2.5 family, optimized for ultra-low latency and cost efficiency. It offers improved throughput, faster token generation, and better performance

Context

1M

In EGP / 1M

5.73

Out EGP / 1M

22.92

gemini-2.5-flash-lite

Gemini 2.5 Pro is Google’s state-of-the-art AI model designed for advanced reasoning, coding, mathematics, and scientific tasks. It employs “thinking” capabilities, enabling it to reason through responses with enhanced accuracy

Context

1M

In EGP / 1M

71.63

Out EGP / 1M

573.02

gemini-2.5-pro-preview

Gemini 3 Flash Preview is a high speed, high value thinking model designed for agentic workflows, multi turn chat, and coding assistance. It delivers near Pro level reasoning and tool

Context

1M

In EGP / 1M

28.65

Out EGP / 1M

171.90

gemini-3-flash-preview
ChatProprietary

Gemini 3.1 Flash Lite is Google’s GA high-efficiency multimodal model optimized for low-latency, high-volume workloads. It supports text, image, video, audio, and PDF inputs, and is designed for lightweight agentic

Context

1M

In EGP / 1M

14.33

Out EGP / 1M

85.95

gemini-3.1-flash-lite

Gemini 3.1 Flash Lite Preview is Google's high-efficiency model optimized for high-volume use cases. It outperforms Gemini 2.5 Flash Lite on overall quality and approaches Gemini 2.5 Flash performance across

Context

1M

In EGP / 1M

14.33

Out EGP / 1M

85.95

gemini-3.1-flash-lite-preview

Gemini 3.1 Pro Preview is Google’s frontier reasoning model, delivering enhanced software engineering performance, improved agentic reliability, and more efficient token usage across complex workflows. Building on the multimodal foundation

Context

1M

In EGP / 1M

114.60

Out EGP / 1M

687.62

gemini-3.1-pro-preview

Gemini 3.1 Pro Preview Custom Tools is a variant of Gemini 3.1 Pro that improves tool selection behavior by preventing overuse of a general bash tool when more efficient third-party

Context

1M

In EGP / 1M

114.60

Out EGP / 1M

687.62

gemini-3.1-pro-preview-customtools
ChatProprietary

Gemini 3.5 Flash is Google's high-efficiency multimodal model, bringing near-Pro level coding and reasoning at Flash-tier cost and speed. It is highly optimized for coding proficiency and parallel agentic execution

Context

1M

In EGP / 1M

85.95

Out EGP / 1M

515.71

gemini-3.5-flash

Gemini 3.5 Flash Lite is a high-efficiency model from Google with upgraded agentic capabilities. It is suited for subagents that execute focused tasks within complex, multi-agent workflows.

Context

1M

In EGP / 1M

17.19

Out EGP / 1M

143.25

gemini-3.5-flash-lite

Gemini 3.6 Flash is a high-efficiency model from Google for coding, agentic workflows, and web and app development. It is designed to produce polished outputs with fewer unnecessary edits and

Context

1M

In EGP / 1M

42.98

Out EGP / 1M

214.88

gemini-3.6-flash
ChatProprietary

Gemini 3.7 Flash is a multimodal model from Google for fast agentic workflows, coding, and complex multi-step reasoning. It is designed for tasks that require responsive performance and reliable multi-step

Context

1M

In EGP / 1M

42.98

Out EGP / 1M

214.88

gemini-3.7-flash

Gemini 3.8 Flash is Google's most intelligent Flash model with significant gains from 3.7 Flash across software engineering, agentic tasks, and multi-step reasoning.

Context

1M

In EGP / 1M

42.98

Out EGP / 1M

214.88

gemini-3.8-flash

This model always redirects to the latest model in the Gemini Flash family.

Context

1M

In EGP / 1M

42.98

Out EGP / 1M

214.88

gemini-flash-latest

This model always redirects to the latest model in the Gemini Pro family.

Context

1M

In EGP / 1M

114.60

Out EGP / 1M

687.62

gemini-pro-latest
Chat

Gemma 2 27B by Google is an open model built from the same research and technology used to create the Gemini models. Gemma models are well-suited for a variety of

Context

8K

In EGP / 1M

37.25

Out EGP / 1M

37.25

gemma-2-27b-it
ChatOpen weights

Gemma 4 26B A4B IT is an instruction-tuned Mixture-of-Experts (MoE) model from Google DeepMind. Despite 25.2B total parameters, only 3.8B activate per token during inference — delivering near-31B quality at

Context

262K

In EGP / 1M

4.01

Out EGP / 1M

19.48

gemma-4-26b-a4b-it
ChatOpen weights

Gemma 4 31B Instruct is Google DeepMind's 30.7B dense multimodal model supporting text and image input with text output. Features a 256K token context window, configurable thinking/reasoning mode, native function

Context

262K

In EGP / 1M

7.45

Out EGP / 1M

21.77

gemma-4-31b-it

Gemini 2.5 Flash Image, a.k.a. "Nano Banana," is now generally available. It is a state of the art image generation model with contextual understanding. It is capable of image generation,

Context

33K

In EGP / 1M

17.19

Out EGP / 1M

143.25

gemini-2.5-flash-image

Gemini 3.1 Flash Image Preview, a.k.a. "Nano Banana 2," is Google’s latest state of the art image generation and editing model, delivering Pro-level visual quality at Flash speed. It combines

Context

66K

In EGP / 1M

28.65

Out EGP / 1M

171.90

gemini-3.1-flash-image-preview

Gemini 3.1 Flash Image, a.k.a. "Nano Banana 2," is Google’s latest state of the art image generation and editing model, delivering Pro-level visual quality at Flash speed. It combines advanced

Context

131K

In EGP / 1M

28.65

Out EGP / 1M

171.90

gemini-3.1-flash-image

Nano Banana 2 Lite (Gemini 3.1 Flash Lite Image) is Google's fastest, most cost-efficient Gemini image model, built for high-velocity developer pipelines and rapid-fire visual exploration. It delivers text-to-image generation

Context

66K

In EGP / 1M

14.33

Out EGP / 1M

85.95

gemini-3.1-flash-lite-image

Nano Banana Pro is Google’s most advanced image-generation and editing model, built on Gemini 3 Pro. It extends the original Nano Banana with significantly improved multimodal reasoning, real-world grounding, and

Context

66K

In EGP / 1M

114.60

Out EGP / 1M

687.62

gemini-3-pro-image-preview

Nano Banana Pro is Google’s most advanced image-generation and editing model, built on Gemini 3 Pro. It extends the original Nano Banana with significantly improved multimodal reasoning, real-world grounding, and

Context

—

Price

7.70 EGP per image

gemini-3-pro-image
OpenAI
ChatOpen weights

gpt-oss-120b is an open-weight, 117B-parameter Mixture-of-Experts (MoE) language model from OpenAI designed for high-reasoning, agentic, and general-purpose production use cases. The model supports configurable reasoning depth, full chain-of-thought access, and native tool use, including function calling, browsing, and structured output generation.

Context

131K

In EGP / 1M

2.12

Out EGP / 1M

9.74

ChatOpen weights

Context

131K

In EGP / 1M

8.60

Out EGP / 1M

34.38

ChatProprietary

Ultra speed version of gpt-oss-120b

Context

131K

In EGP / 1M

11.46

Out EGP / 1M

54.44

OpenAI
ChatOpen weights

gpt-oss-20b is an open-weight 21B parameter model released by OpenAI under the Apache 2.0 license. It uses a Mixture-of-Experts (MoE) architecture with 3.6B active parameters per forward pass, optimized for lower-latency inference. The model is trained in OpenAI’s Harmony response format and supports reasoning level configuration, fine-tuning, and agentic capabilities including function calling, tool use, and structured outputs.

Context

131K

In EGP / 1M

1.72

Out EGP / 1M

8.02

ibm-granite
ChatOpen weights

Granite-4.2-30B is the flagship reasoning model in the Granite 4.2 family. It delivers the strongest performance across reasoning-intensive tasks by leveraging built-in <think>...</think> chain-of-thought. It supports flexible thinking modes — full thinking (default), non-thinking, and low-effort — allowing users to balance depth vs. latency on a per-query basis.

Context

131K

In EGP / 1M

9.17

Out EGP / 1M

37.25

ibm-granite
ChatOpen weights

Granite-4.2-3B is the compact reasoning model in the Granite 4.2 family. Despite its small parameter count, it delivers strong performance on reasoning-intensive tasks by leveraging built-in <think>...</think> chain-of-thought. It supports flexible thinking modes — full thinking (default), non-thinking, and low-effort — allowing users to balance depth vs. latency on a per-query basis.

Context

131K

In EGP / 1M

1.72

Out EGP / 1M

6.88

thenlper
EmbeddingOpen weights

The GTE models are trained by Alibaba DAMO Academy. They are mainly based on the BERT framework and currently offer three different sizes of models, including GTE-large, GTE-base, and GTE-small. The GTE models are trained on a large-scale corpus of relevance text pairs, covering a wide range of domains and scenarios. This enables the GTE models to be applied to various downstream tasks of text embeddings, including information retrieval, semantic textual similarity, text reranking, etc.

Context

512

In EGP / 1M

0.29

Out EGP / 1M

—

thenlper
EmbeddingOpen weights

The GTE models are trained by Alibaba DAMO Academy. They are mainly based on the BERT framework and currently offer three different sizes of models, including GTE-large, GTE-base, and GTE-small. The GTE models are trained on a large-scale corpus of relevance text pairs, covering a wide range of domains and scenarios. This enables the GTE models to be applied to various downstream tasks of text embeddings, including information retrieval, semantic textual similarity, text reranking, etc.

Context

512

In EGP / 1M

0.57

Out EGP / 1M

—

NousResearch
ChatOpen weights

Hermes 3 is a cutting-edge language model that offers advanced capabilities in roleplaying, reasoning, and conversation. It's a fine-tuned version of the Llama-3.1 405B foundation model, designed to align with user needs and provide powerful control. Key features include reliable function calling, structured output, generalist assistant capabilities, and improved code generation. Hermes 3 is competitive with Llama-3.1 Instruct models, with its own strengths and weaknesses.

Context

131K

In EGP / 1M

57.30

Out EGP / 1M

57.30

NousResearch
ChatOpen weights

Hermes 3 is a generalist language model with many improvements over Hermes 2, including advanced agentic capabilities, much better roleplaying, reasoning, multi-turn conversation, long context coherence, and improvements across the board.

Context

131K

In EGP / 1M

40.11

Out EGP / 1M

40.11

ibm-granite
Chat

Granite-4.0-H-Micro is a 3B parameter from the Granite 4 family of models. These models are the latest in a series of models released by IBM. They are fine-tuned for long

Context

131K

In EGP / 1M

0.97

Out EGP / 1M

6.42

granite-4.0-h-micro
ibm-granite
ChatOpen weights

Granite 4.2 8B is a dense reasoning model from IBM. It is suited for mathematics, code generation, multilingual dialogue, and agentic workflows that need multi-step reasoning. It supports full, low-effort,

Context

131K

In EGP / 1M

3.44

Out EGP / 1M

14.33

granite-4.2-8b
G42 / Inception
Chat

Mercury 2 is an extremely fast reasoning LLM, and the first reasoning diffusion LLM (dLLM). Instead of generating tokens sequentially, Mercury 2 produces and refines multiple tokens in parallel, achieving

Context

128K

In EGP / 1M

14.33

Out EGP / 1M

42.98

mercury-2
G42 / Inception
Chat

Mercury 2.5 is the fastest reasoning LLM, and the latest diffusion LLM (dLLM) from Inception. Instead of generating tokens sequentially, Mercury 2.5 produces and refines multiple tokens in parallel, achieving

Context

260K

In EGP / 1M

2.29

Out EGP / 1M

8.60

mercury-2.5
ChatOpen weights

*Ling-3.0-flash* is a *124B-parameter Mixture-of-Experts (MoE) model*, with approximately *5.1B parameters activated per token*. The model is designed with *token efficiency and production-scale agentic inference* as key priorities, enabling developers

Context

131K

In EGP / 1M

3.44

Out EGP / 1M

10.31

ling-3.0-flash
ChatOpen weights

Ling 3.0 Flash Fin is a finance-focused mixture-of-experts model from InclusionAI, built on Ling 3.0 Flash with 5.1B active parameters out of 124B total. It is designed for real-world investment

Context

262K

In EGP / 1M

3.44

Out EGP / 1M

10.31

ling-3.0-flash-fin
ChatOpen weights

Ling 3.0 Flash VL builds on Ling 3.0 Flash (124B total / 5.5B active MoE from InclusionAI), further strengthening its language capabilities while adding native visual perception and advanced visual

Context

131K

In EGP / 1M

3.44

Out EGP / 1M

10.31

ling-3.0-flash-vl
Chat

Schematron V2 Small is a 3B-parameter HTML-to-JSON extraction model from Inference.net. It prioritizes extraction quality for complex schemas and long pages. Extraction instructions must be supplied through a JSON schema

Context

128K

In EGP / 1M

2.87

Out EGP / 1M

13.18

schematron-v2-small
Chat

Schematron V2 Turbo is a 3B-parameter HTML-to-JSON extraction model from Inference.net. It prioritizes throughput for high-volume extraction workloads. Extraction instructions must be supplied through a JSON schema in response_format rather

Context

128K

In EGP / 1M

1.72

Out EGP / 1M

8.60

schematron-v2-turbo
jaredpalmer
Decision

Kev 4B is a small open-weight decision model from Jared Palmer, built as a LoRA adapter and pointer head on Qwen3.5-4B-Base and served over the same /v1/systemone contract as TypeSafe's

Context

8K

In EGP / 1M

2.41

Out EGP / 1M

Free

kev-4b
ChatOpen weights

Llama 3.3-70B Turbo is a highly optimized version of the Llama 3.3-70B model, utilizing FP8 quantization to deliver significantly faster inference speeds with a minor trade-off in accuracy. The model is designed to be helpful, safe, and flexible, with a focus on responsible deployment and mitigating potential risks such as bias, toxicity, and misinformation. It achieves state-of-the-art performance on various benchmarks, including conversational tasks, language translation, and text generation.

Context

66K

In EGP / 1M

5.73

Out EGP / 1M

18.34

ChatOpen weights

The Llama 4 collection of models are natively multimodal AI models that enable text and multimodal experiences. These models leverage a mixture-of-experts architecture to offer industry-leading performance in text and image understanding. Llama 4 Scout, a 17 billion parameter model with 16 experts

Context

328K

In EGP / 1M

5.73

Out EGP / 1M

17.19

ChatOpen weights

Llama Guard 4 is a natively multimodal safety classifier with 12 billion parameters trained jointly on text and multiple images. Llama Guard 4 is a dense architecture pruned from the Llama 4 Scout pre-trained model and fine-tuned for content safety classification. Similar to previous versions, it can be used to classify content in both LLM inputs (prompt classification) and in LLM responses (response classification). It itself acts as an LLM: it generates text in its output that indicates whether a given prompt or response is safe or unsafe, and if unsafe, it also lists the content categories violated.

Context

164K

In EGP / 1M

10.31

Out EGP / 1M

10.31

EmbeddingOpen weights

The llama-nemotron-embed-vl-1b-v2 is a high-performance multimodal embedding model designed to transform text queries and document images into dense vector representations for advanced retrieval systems. It excels at understanding complex visual content like charts, tables, and infographics.

Context

10K

In EGP / 1M

0.57

Out EGP / 1M

—

Chat

An attempt to recreate Claude-style verbosity, but don't expect the same level of coherence or memory. Meant for use in roleplay/narrative situations.

Context

8K

In EGP / 1M

22.92

Out EGP / 1M

42.98

weaver
Chat

LongCat 2.0 is a sparse mixture-of-experts language model from Meituan, with 48B active parameters out of 1.6T total. It is suited for coding, repository-level changes, long-horizon problem solving, and agentic

Context

1M

In EGP / 1M

17.19

Out EGP / 1M

68.76

longcat-2.0
ChatOpen weights

Meta developed and released the Meta Llama 3.1 family of large language models (LLMs), a collection of pretrained and instruction tuned generative text models in 8B, 70B and 405B sizes

Context

131K

In EGP / 1M

22.92

Out EGP / 1M

22.92

ChatOpen weights

Meta developed and released the Meta Llama 3.1 family of large language models (LLMs), a collection of pretrained and instruction tuned generative text models in 8B, 70B and 405B sizes

Context

131K

In EGP / 1M

1.15

Out EGP / 1M

2.29

Meta's latest class of model (Llama 3.1) launched with a variety of sizes & flavors. This 70B instruct-tuned version is optimized for high quality dialogue usecases. It has demonstrated strong

Context

131K

In EGP / 1M

22.92

Out EGP / 1M

22.92

llama-3.1-70b-instruct

Meta's latest class of model (Llama 3.1) launched with a variety of sizes & flavors. This 8B instruct-tuned version is fast and efficient. It has demonstrated strong performance compared to

Context

131K

In EGP / 1M

2.87

Out EGP / 1M

4.58

llama-3.1-8b-instruct

Llama 3.2 1B is a 1-billion-parameter language model focused on efficiently performing natural language tasks, such as summarization, dialogue, and multilingual text analysis. Its smaller size allows it to operate

Context

60K

In EGP / 1M

1.55

Out EGP / 1M

11.52

llama-3.2-1b-instruct

Llama 3.2 3B is a 3-billion-parameter multilingual large language model, optimized for advanced natural language processing tasks like dialogue generation, reasoning, and summarization. Designed with the latest transformer architecture, it

Context

131K

In EGP / 1M

2.87

Out EGP / 1M

18.91

llama-3.2-3b-instruct

The Meta Llama 3.3 multilingual large language model (LLM) is a pretrained and instruction tuned generative model in 70B (text in/text out). The Llama 3.3 instruction tuned text only model

Context

66K

In EGP / 1M

5.73

Out EGP / 1M

18.34

llama-3.3-70b-instruct

Llama 4 Maverick 17B Instruct (128E) is a high-capacity multimodal language model from Meta, built on a mixture-of-experts (MoE) architecture with 128 experts and 17 billion active parameters per forward

Context

128K

In EGP / 1M

10.74

Out EGP / 1M

37.39

llama-4-maverick

Llama 4 Scout 17B Instruct (16E) is a mixture-of-experts (MoE) language model developed by Meta, activating 17 billion parameters out of a total of 109B. It supports native multimodal input

Context

328K

In EGP / 1M

5.73

Out EGP / 1M

17.19

llama-4-scout
ChatOpen weights

Muse Glimmer 30B is a dense, open-weight multimodal model from Meta Superintelligence Labs, distilled from Muse Spark and optimized for autonomous agents on consumer hardware. It is suited for long-horizon

Context

131K

In EGP / 1M

17.19

Out EGP / 1M

68.76

muse-glimmer-30b
inclusionAI
imageOpen weights

Design-native, open-weight text-to-image (6B) — ranked #1 among open-weight models on Artificial Analysis's UI/UX Design leaderboard. Renders legible UI and poster text plus cohesive graphics, illustration and photography in ~1.7s. MIT-licensed.

Context

—

Price

0.573 EGP per 1024p image

imageOpen weights

Open-weight layer decomposition (6B) from the Ming-Image 0.1 Design family — splits a flattened design or composite into stacked, transparent, editable PNG layers (text, subject, background) for localizable, editable design workflows. MIT-licensed.

Context

—

Price

0.573 EGP per 1024p image

Chat

MiniMax-M1 is a large-scale, open-weight reasoning model designed for extended context and high-efficiency inference. It leverages a hybrid Mixture-of-Experts (MoE) architecture paired with a custom "lightning attention" mechanism, allowing it

Context

1M

In EGP / 1M

31.52

Out EGP / 1M

126.06

minimax-m1
Chat

MiniMax-M2 is a compact, high-efficiency large language model optimized for end-to-end coding and agentic workflows. With 10 billion activated parameters (230 billion total), it delivers near-frontier intelligence across general reasoning,

Context

197K

In EGP / 1M

17.19

Out EGP / 1M

68.76

minimax-m2

MiniMax M2-her is a dialogue-first large language model built for immersive roleplay, character-driven chat, and expressive multi-turn conversations. Designed to stay consistent in tone and personality, it supports rich message

Context

66K

In EGP / 1M

17.19

Out EGP / 1M

68.76

minimax-m2-her
Chat

MiniMax-M2.1 is a lightweight, state-of-the-art large language model optimized for coding, agentic workflows, and modern application development. With only 10 billion activated parameters, it delivers a major jump in real-world

Context

205K

In EGP / 1M

17.19

Out EGP / 1M

68.76

minimax-m2.1
Chat

MiniMax-M2.5 is a SOTA large language model designed for real-world productivity. Trained in a diverse range of complex real-world digital working environments, M2.5 builds upon the coding expertise of M2.1

Context

200K

In EGP / 1M

15.47

Out EGP / 1M

61.89

minimax-m2.5
Chat

MiniMax-M2.7 is a next-generation large language model designed for autonomous, real-world productivity and continuous improvement. Built to actively participate in its own evolution, M2.7 integrates advanced agentic capabilities through multi-agent

Context

197K

In EGP / 1M

12.03

Out EGP / 1M

48.13

minimax-m2.7
ChatOpen weights

MiniMax-M3 is a multimodal foundation model from MiniMax. It supports text, image, and video inputs with text output, a 1M-token context window, and is suited for long-horizon agentic work, coding,

Context

524K

In EGP / 1M

16.04

Out EGP / 1M

63.03

minimax-m3
Chat

MiniMax-01 is a combines MiniMax-Text-01 for text generation and MiniMax-VL-01 for image understanding. It has 456 billion parameters, with 45.9 billion parameters activated per inference, and can handle a context

Context

1M

In EGP / 1M

11.46

Out EGP / 1M

63.03

minimax-01
Mistral AI
Chat

This is Mistral AI's flagship model, Mistral Large 2 (version `mistral-large-2407`). It's a proprietary weights-available model and excels at reasoning, code, JSON, chat, and more. Read the launch announcement here

Context

128K

In EGP / 1M

114.60

Out EGP / 1M

343.81

mistral-large
Mistral AI
Chat

This is Mistral AI's flagship model, Mistral Large 2 (version mistral-large-2407). It's a proprietary weights-available model and excels at reasoning, code, JSON, chat, and more. Read the launch announcement here

Context

131K

In EGP / 1M

114.60

Out EGP / 1M

343.81

mistral-large-2407
ChatOpen weights

12B model trained jointly by Mistral AI and NVIDIA, it significantly outperforms existing models smaller or similar in size.

Context

131K

In EGP / 1M

1.09

Out EGP / 1M

1.72

ChatOpen weights

Mistral Small 3 is a 24B-parameter language model optimized for low-latency performance across common AI tasks. Released under the Apache 2.0 license, it features both pre-trained and instruction-tuned versions designed for efficient local deployment. The model achieves 81% accuracy on the MMLU benchmark and performs competitively with larger models like Llama 3.3 70B and Qwen 32B, while operating at three times the speed on equivalent hardware.

Context

33K

In EGP / 1M

2.87

Out EGP / 1M

4.58

ChatOpen weights

Mistral-Small-3.2-24B-Instruct is a drop-in upgrade over the 3.1 release, with markedly better instruction following, roughly half the infinite-generation errors, and a more robust function-calling interface—while otherwise matching or slightly improving on all previous text and vision benchmarks.

Context

128K

In EGP / 1M

4.30

Out EGP / 1M

11.46

Chat

Mistral's cutting-edge language model for coding released end of July 2025. Codestral specializes in low-latency, high-frequency tasks such as fill-in-the-middle (FIM), code correction and test generation. Blog Post

Context

256K

In EGP / 1M

17.19

Out EGP / 1M

51.57

codestral-2508
Chat

Devstral 2 is a state-of-the-art open-source model by Mistral AI specializing in agentic coding. It is a 123B-parameter dense transformer model supporting a 256K context window. Devstral 2 supports exploring

Context

262K

In EGP / 1M

22.92

Out EGP / 1M

114.60

devstral-2512

The largest model in the Ministral 3 family, Ministral 3 14B offers frontier capabilities and performance comparable to its larger Mistral Small 3.2 24B counterpart. A powerful and efficient language

Context

262K

In EGP / 1M

11.46

Out EGP / 1M

11.46

ministral-14b-2512

The smallest model in the Ministral 3 family, Ministral 3 3B is a powerful, efficient tiny language model with vision capabilities.

Context

131K

In EGP / 1M

5.73

Out EGP / 1M

5.73

ministral-3b-2512

A balanced model in the Ministral 3 family, Ministral 3 8B is a powerful, efficient tiny language model with vision capabilities.

Context

262K

In EGP / 1M

8.60

Out EGP / 1M

8.60

ministral-8b-2512
Chat

Mistral Medium 3 is a high-performance enterprise-grade language model designed to deliver frontier-level capabilities at significantly reduced operational cost. It balances state-of-the-art reasoning and multimodal performance with 8× lower cost

Context

131K

In EGP / 1M

22.92

Out EGP / 1M

114.60

mistral-medium-3
Chat

Mistral Medium 3.1 is an updated version of Mistral Medium 3, which is a high-performance enterprise-grade language model designed to deliver frontier-level capabilities at significantly reduced operational cost. It balances

Context

131K

In EGP / 1M

22.92

Out EGP / 1M

114.60

mistral-medium-3.1
Chat

Mistral Medium 3.5 is a dense 128B instruction-following model from Mistral AI. It supports text and image inputs with text output, and is designed for agentic workflows, coding, and complex

Context

262K

In EGP / 1M

85.95

Out EGP / 1M

429.76

mistral-medium-3-5
Mistral AI
Chat

A 12B parameter model with a 128k token context length built by Mistral in collaboration with NVIDIA. The model is multilingual, supporting English, French, German, Spanish, Italian, Portuguese, Chinese, Japanese,

Context

131K

In EGP / 1M

1.09

Out EGP / 1M

1.72

mistral-nemo

Mistral Small 3.1 24B Instruct is an upgraded variant of Mistral Small 3 (2501), featuring 24 billion parameters with advanced multimodal capabilities. It provides state-of-the-art performance in text-based reasoning and

Context

128K

In EGP / 1M

20.11

Out EGP / 1M

31.80

mistral-small-3.1-24b-instruct

Mistral-Small-3.2-24B-Instruct-2506 is an updated 24B parameter model from Mistral optimized for instruction following, repetition reduction, and improved function calling. Compared to the 3.1 release, version 3.2 significantly improves accuracy on

Context

256K

In EGP / 1M

5.37

Out EGP / 1M

14.33

mistral-small-3.2-24b-instruct
Chat

Kimi K2 Instruct is a large-scale Mixture-of-Experts (MoE) language model developed by Moonshot AI, featuring 1 trillion total parameters with 32 billion active per forward pass. It is optimized for

Context

131K

In EGP / 1M

32.66

Out EGP / 1M

131.79

kimi-k2
Chat

Kimi K2 0905 is the September update of Kimi K2 0711. It is a large-scale Mixture-of-Experts (MoE) language model developed by Moonshot AI, featuring 1 trillion total parameters with 32

Context

262K

In EGP / 1M

34.38

Out EGP / 1M

143.25

kimi-k2-0905

Kimi K2 Thinking is Moonshot AI’s most advanced open reasoning model to date, extending the K2 series into agentic, long-horizon reasoning. Built on the trillion-parameter Mixture-of-Experts (MoE) architecture introduced in

Context

262K

In EGP / 1M

34.38

Out EGP / 1M

143.25

kimi-k2-thinking
moonshotai
Chat

Kimi K2.5 is Moonshot AI's native multimodal model, delivering state-of-the-art visual coding capability and a self-directed agent swarm paradigm. Built on Kimi K2 with continued pretraining over approximately 15T mixed

Context

262K

In EGP / 1M

25.79

Out EGP / 1M

128.93

kimi-k2.5
moonshotai
ChatOpen weights

Kimi K2.6 is Moonshot AI's next-generation multimodal model, designed for long-horizon coding, coding-driven UI/UX generation, and multi-agent orchestration. It handles complex end-to-end coding tasks across Python, Rust, and Go, and

Context

262K

In EGP / 1M

42.98

Out EGP / 1M

200.56

kimi-k2.6
Chat

MoonshotAI: Kimi K2.7 Code is a coding-focused model in Moonshot AI's Kimi K2 family, built to complete end-to-end programming tasks reliably over long contexts. It uses a native multimodal mixture-of-experts

Context

262K

In EGP / 1M

38.46

Out EGP / 1M

191.96

kimi-k2.7-code
moonshotai
ChatOpen weights

Kimi K3 is a 2.8T parameter open-weight multimodal reasoning model from Moonshot AI. It is suited for complex coding, knowledge work, and long-horizon agentic workflows, and is particularly strong at

Context

1M

In EGP / 1M

163.31

Out EGP / 1M

816.55

kimi-k3
~moonshotai
Chat

This model always redirects to the latest model in the Kimi family.

Context

1M

In EGP / 1M

16.04

Out EGP / 1M

573.02

kimi-latest
Chat

Morph's fastest apply model for code edits. ~10,500 tokens/sec with 96% accuracy for rapid code transformations. The model requires the prompt to be in the following format: <instruction>{instruction}</instruction> <code>{initial_code}</code> <update>{edit_snippet}</update>

Context

82K

In EGP / 1M

45.84

Out EGP / 1M

68.76

morph-v3-fast
Chat

Morph's high-accuracy apply model for complex code edits. ~4,500 tokens/sec with 98% accuracy for precise code transformations. The model requires the prompt to be in the following format: <instruction>{instruction}</instruction> <code>{initial_code}</code>

Context

262K

In EGP / 1M

51.57

Out EGP / 1M

108.87

morph-v3-large
EmbeddingOpen weights

We present a sentence transformation model that maps sentences and paragraphs to a 768-dimensional dense vector space, suitable for semantic search tasks. The model is trained on 215 million question-answer pairs from various sources, including WikiAnswers, PAQ, Stack Exchange, MS MARCO, GOOAQ, Amazon QA, Yahoo Answers, Search QA, ELI5, and Natural Questions. Our model uses a contrastive learning objective.

Context

512

In EGP / 1M

0.29

Out EGP / 1M

—

EmbeddingOpen weights

The Multilingual-E5-large model is a 24-layer text embedding model with an embedding size of 1024, trained on a mixture of multilingual datasets and supporting 100 languages.

Context

512

In EGP / 1M

0.57

Out EGP / 1M

—

EmbeddingOpen weights

The Multilingual-E5 models, initialized from XLM-RoBERTa, support up to 512 tokens per input — any longer text will be silently truncated. To ensure optimal performance, always prefix inputs with “query:” or “passage:”, as the model was explicitly trained with this format.

Context

512

In EGP / 1M

0.57

Out EGP / 1M

—

Google
imageProprietary

Nano Banana 2 makes high quality image generation and editing mainstream. The model also introduces Grounding with Google Image Search to enable long-tail entity recognition and enhance visual understanding.

Context

—

Price

3.85 EGP per image

imageProprietary

Nano Banana 2 Lite enables high speed image generation a reality, while keeping the high quality editing of Nano Banana.

Context

—

Price

1.93 EGP per image

Google
imageProprietary

Nano Banana Pro is designed to tackle the most challenging image generation by incorporating state-of-the-art reasoning capabilities. It is the best model for complex and multi-turn image generation and editing.

Context

—

Price

7.70 EGP per image

ChatOpen weights

NVIDIA Nemotron 3 Nano is an open small reasoning model optimized for fast, cost-efficient inference in agentic and production workloads. Built with a hybrid Mixture-of-Experts (MoE) and Mamba-Transformer architecture, it delivers strong multi-step reasoning, high token throughput, stable latency with predictable cost, and efficient deployment for agent-based systems. Designed for real-world AI systems where reasoning can generate significantly more tokens per prompt, Nemotron Nano reduces compute cost while maintaining strong reasoning quality.

Context

262K

In EGP / 1M

2.87

Out EGP / 1M

11.46

ChatOpen weights

Nemotron Content Safety 3.5 is a multimodal safety classifier developed by NVIDIA. A compact safety model that handles text, images, and custom policies. It outputs a safe/unsafe classification plus a reasoning trace, and can be used as an inference-time guardrail, as a judge for LLM safety testing and evaluation, or with the accompanying training dataset to post-train models for safer behavior.

Context

131K

In EGP / 1M

11.46

Out EGP / 1M

11.46

Chat

Nex-N2.5 is an agentic model built to turn goals into working, verified outcomes. Its core strength is agentic coding within a visual feedback loop: it can explore codebases, implement multi-file

Context

262K

In EGP / 1M

1.43

Out EGP / 1M

5.73

nex-n2.5-mini
Chat

Nex-N2.5 is an agentic model built to turn goals into working, verified outcomes. Its core strength is agentic coding within a visual feedback loop: it can explore codebases, implement multi-file

Context

262K

In EGP / 1M

4.30

Out EGP / 1M

14.33

nex-n2.5-pro
nousresearch
Chat

Hermes 4 is a large-scale reasoning model built on Meta-Llama-3.1-405B and released by Nous Research. It introduces a hybrid reasoning mode, where the model can choose to deliberate internally with

Context

131K

In EGP / 1M

57.30

Out EGP / 1M

171.90

hermes-4-405b
ChatOpen weights

NVIDIA Nemotron 3 Super is a hybrid Mixture-of-Experts (MoE) model engineered for highest compute efficiency and accuracy in multi-agent applications and specialized agentic systems. It is optimized to run many collaborating agents per application on a single GPU, delivering high accuracy for reasoning, tool use, and instruction following.

Context

262K

In EGP / 1M

4.87

Out EGP / 1M

22.92

ChatOpen weights

Nemotron 3 Ultra is built for, frontier reasoning, orchestration, coding agents, deep research, and complex enterprise workflows. It delivers up to 5x faster inference and up to 30% lower cost for agentic workloads while supporting up to 1M token context.

Context

262K

In EGP / 1M

28.65

Out EGP / 1M

126.06

ChatOpen weights

NVIDIA Nemotron 3.5 Lightning is NVIDIA's fastest open model for always-on agents and high-volume specialized tasks. It delivers a substantial leap in agentic capability over its predecessor Nemotron 3 Nano, with up to 4x higher throughput on a 1M-token context.

Context

262K

In EGP / 1M

3.44

Out EGP / 1M

9.17

NVIDIA Nemotron 3 Super is a 120B-parameter open hybrid MoE model, activating just 12B parameters for maximum compute efficiency and accuracy in complex multi-agent applications. Built on a hybrid Mamba-Transformer

Context

262K

In EGP / 1M

4.58

Out EGP / 1M

25.79

nemotron-3-super-120b-a12b

NVIDIA Nemotron 3 Ultra is an open frontier-reasoning and orchestration model from NVIDIA, with 55B active parameters out of 550B total (MoE). Built on a hybrid Transformer-Mamba mixture-of-experts architecture, it

Context

203K

In EGP / 1M

34.38

Out EGP / 1M

137.52

nemotron-3-ultra-550b-a55b

NVIDIA Nemotron 3.5 Content Safety is a compact 4B-parameter multimodal guardrail model from NVIDIA, fine-tuned from Google Gemma-3-4B. It moderates both inputs to and responses from LLMs and VLMs, accepting

Context

131K

In EGP / 1M

11.46

Out EGP / 1M

11.46

nemotron-3.5-content-safety

NVIDIA Nemotron 3.5 Lightning is an open mixture-of-experts model from NVIDIA, with 3B active parameters out of 30B total. It is suited for high-throughput agentic workloads and specialized tasks that

Context

262K

In EGP / 1M

3.41

Out EGP / 1M

9.74

nemotron-3.5-lightning
Chat

This model always redirects to the latest model in the GPT Astra family.

Context

1M

In EGP / 1M

573.02

Out EGP / 1M

2865.08

gpt-astra-latest

GPT Chat Latest points to OpenAI's stable API alias `chat-latest` that always resolves to the latest Instant chat model used in ChatGPT. As OpenAI rolls out new Instant model updates

Context

400K

In EGP / 1M

286.51

Out EGP / 1M

1719.05

gpt-chat-latest
Chat

This model always redirects to the latest model in the GPT Luna family.

Context

1M

In EGP / 1M

5.73

Out EGP / 1M

28.65

gpt-luna-latest
Chat

This model always redirects to the latest model in the GPT Mini family.

Context

400K

In EGP / 1M

42.98

Out EGP / 1M

257.86

gpt-mini-latest
Chat

This model always redirects to the latest model in the GPT Sol family.

Context

1M

In EGP / 1M

114.60

Out EGP / 1M

573.02

gpt-sol-latest
Chat

This model always redirects to the latest model in the GPT Terra family.

Context

1M

In EGP / 1M

114.60

Out EGP / 1M

687.62

gpt-terra-latest
Chat

GPT-3.5 Turbo is OpenAI's fastest model. It can understand and generate natural language or code, and is optimized for chat and traditional completion tasks. Training data up to Sep 2021.

Context

16K

In EGP / 1M

28.65

Out EGP / 1M

85.95

gpt-3.5-turbo

GPT-3.5 Turbo is OpenAI's fastest model. It can understand and generate natural language or code, and is optimized for chat and traditional completion tasks. Training data up to Sep 2021.

Context

4K

In EGP / 1M

57.30

Out EGP / 1M

114.60

gpt-3.5-turbo-0613

This model offers four times the context length of gpt-3.5-turbo, allowing it to support approximately 20 pages of text in a single request at a higher cost. Training data: up

Context

16K

In EGP / 1M

171.90

Out EGP / 1M

229.21

gpt-3.5-turbo-16k
OpenAI
Chat

OpenAI's flagship model, GPT-4 is a large-scale multimodal language model capable of solving difficult problems with greater accuracy than previous models due to its broader general knowledge and advanced reasoning

Context

8K

In EGP / 1M

1719.05

Out EGP / 1M

3438.09

gpt-4
Chat

The latest GPT-4 Turbo model with vision capabilities. Vision requests can now use JSON mode and function calling. Training data: up to December 2023.

Context

128K

In EGP / 1M

573.02

Out EGP / 1M

1719.05

gpt-4-turbo
OpenAI
Chat

GPT-4.1 is a flagship large language model optimized for advanced instruction following, real-world software engineering, and long-context reasoning. It supports a 1 million token context window and outperforms GPT-4o and

Context

1M

In EGP / 1M

114.60

Out EGP / 1M

458.41

gpt-4.1
Chat

GPT-4.1 Mini is a mid-sized model delivering performance competitive with GPT-4o at substantially lower latency and cost. It retains a 1 million token context window and scores 45.1% on hard

Context

1M

In EGP / 1M

22.92

Out EGP / 1M

91.68

gpt-4.1-mini
Chat

For tasks that demand low latency, GPT‑4.1 nano is the fastest and cheapest model in the GPT-4.1 series. It delivers exceptional performance at a small size with its 1 million

Context

1M

In EGP / 1M

5.73

Out EGP / 1M

22.92

gpt-4.1-nano
OpenAI
Chat

GPT-4o ("o" for "omni") is OpenAI's latest AI model, supporting both text and image inputs with text outputs. It maintains the intelligence level of GPT-4 Turbo while being twice as

Context

128K

In EGP / 1M

143.25

Out EGP / 1M

573.02

gpt-4o

GPT-4o ("o" for "omni") is OpenAI's latest AI model, supporting both text and image inputs with text outputs. It maintains the intelligence level of GPT-4 Turbo while being twice as

Context

128K

In EGP / 1M

286.51

Out EGP / 1M

859.52

gpt-4o-2024-05-13

The 2024-08-06 version of GPT-4o offers improved performance in structured outputs, with the ability to supply a JSON schema in the respone_format. Read more here. GPT-4o ("o" for "omni") is

Context

128K

In EGP / 1M

143.25

Out EGP / 1M

573.02

gpt-4o-2024-08-06

The 2024-11-20 version of GPT-4o offers a leveled-up creative writing ability with more natural, engaging, and tailored writing to improve relevance & readability. It’s also better at working with uploaded

Context

128K

In EGP / 1M

143.25

Out EGP / 1M

573.02

gpt-4o-2024-11-20
Chat

GPT-4o mini is OpenAI's newest model after GPT-4 Omni, supporting both text and image inputs with text outputs. As their most advanced small model, it is many multiples more affordable

Context

128K

In EGP / 1M

8.60

Out EGP / 1M

34.38

gpt-4o-mini

GPT-4o mini is OpenAI's newest model after GPT-4 Omni, supporting both text and image inputs with text outputs. As their most advanced small model, it is many multiples more affordable

Context

128K

In EGP / 1M

8.60

Out EGP / 1M

34.38

gpt-4o-mini-2024-07-18
OpenAI
Chat

GPT-5 is OpenAI’s most advanced model, offering major improvements in reasoning, code quality, and user experience. It is optimized for complex tasks that require step-by-step reasoning, instruction following, and accuracy

Context

400K

In EGP / 1M

71.63

Out EGP / 1M

573.02

gpt-5
Chat

GPT-5 Image combines OpenAI's GPT-5 model with state-of-the-art image generation capabilities. It offers major improvements in reasoning, code quality, and user experience while incorporating GPT Image 1's superior instruction following,

Context

400K

In EGP / 1M

573.02

Out EGP / 1M

573.02

gpt-5-image

GPT-5 Image Mini combines OpenAI's advanced language capabilities, powered by GPT-5 Mini, with GPT Image 1 Mini for efficient image generation. This natively multimodal model features superior instruction following, text

Context

400K

In EGP / 1M

143.25

Out EGP / 1M

114.60

gpt-5-image-mini
Chat

GPT-5 Mini is a compact version of GPT-5, designed to handle lighter-weight reasoning tasks. It provides the same instruction-following and safety-tuning benefits as GPT-5, but with reduced latency and cost

Context

400K

In EGP / 1M

14.33

Out EGP / 1M

114.60

gpt-5-mini
Chat

GPT-5-Nano is the smallest and fastest variant in the GPT-5 system, optimized for developer tools, rapid interactions, and ultra-low latency environments. While limited in reasoning depth compared to its larger

Context

400K

In EGP / 1M

2.87

Out EGP / 1M

22.92

gpt-5-nano
OpenAI
Chat

GPT-5.1 is the latest frontier-grade model in the GPT-5 series, offering stronger general-purpose reasoning, improved instruction adherence, and a more natural conversational style compared to GPT-5. It uses adaptive reasoning

Context

400K

In EGP / 1M

71.63

Out EGP / 1M

573.02

gpt-5.1
Chat

GPT-5.1-Codex is a specialized version of GPT-5.1 optimized for software engineering and coding workflows. It is designed for both interactive development sessions and long, independent execution of complex engineering tasks

Context

400K

In EGP / 1M

71.63

Out EGP / 1M

573.02

gpt-5.1-codex

GPT-5.1-Codex-Max is OpenAI’s latest agentic coding model, designed for long-running, high-context software development tasks. It is based on an updated version of the 5.1 reasoning stack and trained on agentic

Context

400K

In EGP / 1M

71.63

Out EGP / 1M

573.02

gpt-5.1-codex-max

GPT-5.1-Codex-Mini is a smaller and faster version of GPT-5.1-Codex

Context

400K

In EGP / 1M

14.33

Out EGP / 1M

114.60

gpt-5.1-codex-mini
OpenAI
Chat

GPT-5.2 is the latest frontier-grade model in the GPT-5 series, offering stronger agentic and long context perfomance compared to GPT-5.1. It uses adaptive reasoning to allocate computation dynamically, responding quickly

Context

400K

In EGP / 1M

100.28

Out EGP / 1M

802.22

gpt-5.2
Chat

GPT-5.2 Pro is OpenAI’s most advanced model, offering major improvements in agentic coding and long context performance over GPT-5 Pro. It is optimized for complex tasks that require step-by-step reasoning,

Context

400K

In EGP / 1M

1203.33

Out EGP / 1M

9626.66

gpt-5.2-pro
Chat

GPT-5.2-Codex is an upgraded version of GPT-5.1-Codex optimized for software engineering and coding workflows. It is designed for both interactive development sessions and long, independent execution of complex engineering tasks

Context

400K

In EGP / 1M

100.28

Out EGP / 1M

802.22

gpt-5.2-codex
Chat

GPT-5.3-Codex is OpenAI’s most advanced agentic coding model, combining the frontier software engineering performance of GPT-5.2-Codex with the broader reasoning and professional knowledge capabilities of GPT-5.2. It achieves state-of-the-art results

Context

400K

In EGP / 1M

100.28

Out EGP / 1M

802.22

gpt-5.3-codex
OpenAI
Chat

GPT-5.4 is OpenAI’s latest frontier model, unifying the Codex and GPT lines into a single system. It features a 1M+ token context window (922K input, 128K output) with support for

Context

1M

In EGP / 1M

143.25

Out EGP / 1M

859.52

gpt-5.4

GPT-5.4 Image 2 combines OpenAI's GPT-5.4 model with state-of-the-art image generation capabilities from GPT Image 2. It enables rich multimodal workflows, allowing users to seamlessly move between reasoning, coding, and

Context

272K

In EGP / 1M

458.41

Out EGP / 1M

859.52

gpt-5.4-image-2
Chat

GPT-5.4 mini brings the core capabilities of GPT-5.4 to a faster, more efficient model optimized for high-throughput workloads. It supports text and image inputs with strong performance across reasoning, coding,

Context

400K

In EGP / 1M

42.98

Out EGP / 1M

257.86

gpt-5.4-mini
Chat

GPT-5.4 nano is the most lightweight and cost-efficient variant of the GPT-5.4 family, optimized for speed-critical and high-volume tasks. It supports text and image inputs and is designed for low-latency

Context

400K

In EGP / 1M

11.46

Out EGP / 1M

71.63

gpt-5.4-nano
Chat

GPT-5.4 Pro is OpenAI's most advanced model, building on GPT-5.4's unified architecture with enhanced reasoning capabilities for complex, high-stakes tasks. It features a 1M+ token context window (922K input, 128K

Context

1M

In EGP / 1M

1719.05

Out EGP / 1M

10314.28

gpt-5.4-pro
OpenAI
Chat

GPT-5.5 is OpenAI’s frontier model designed for complex professional workloads, building on GPT-5.4 with stronger reasoning, higher reliability, and improved token efficiency on hard tasks. It features a 1M+ token

Context

1M

In EGP / 1M

286.51

Out EGP / 1M

1719.05

gpt-5.5
Chat

GPT-5.5 Pro is OpenAI’s high-capability model optimized for deep reasoning and accuracy on complex, high-stakes workloads. It features a 1M+ token context window (922K input, 128K output) with support for

Context

1M

In EGP / 1M

1719.05

Out EGP / 1M

10314.28

gpt-5.5-pro
Chat

GPT-5.6 Luna is a fast, cost-efficient model in OpenAI's GPT-5.6 series. It is suited for high-volume, latency-sensitive tasks such as chat, classification, and lightweight agentic workflows, providing capable reasoning for

Context

1M

In EGP / 1M

11.46

Out EGP / 1M

68.76

gpt-5.6-luna

GPT-5.6 Luna Pro is the same underlying model as GPT-5.6 Luna, served with `reasoning.mode` set to `pro` for higher-quality responses on complex tasks. Learn more in OpenAI's docs: https://developers.openai.com/api/docs/guides/reasoning#reasoning-mode

Context

1M

In EGP / 1M

11.46

Out EGP / 1M

68.76

gpt-5.6-luna-pro
Chat

GPT-5.6 Sol is the flagship model in OpenAI's GPT-5.6 series. It is suited for complex reasoning, coding, and agentic workflows, and is particularly strong at command-line and multi-step coding tasks

Context

1M

In EGP / 1M

114.60

Out EGP / 1M

573.02

gpt-5.6-sol

GPT-5.6 Sol Pro is the same underlying model as GPT-5.6 Sol, served with `reasoning.mode` set to `pro` for higher-quality responses on complex tasks. Learn more in OpenAI's docs: https://developers.openai.com/api/docs/guides/reasoning#reasoning-mode

Context

1M

In EGP / 1M

229.21

Out EGP / 1M

1146.03

gpt-5.6-sol-pro
Chat

GPT-5.6 Terra is a balanced model in OpenAI's GPT-5.6 series, positioned between the flagship Sol tier and the cost-efficient Luna tier. It is suited for everyday coding, reasoning, and agentic

Context

1M

In EGP / 1M

114.60

Out EGP / 1M

687.62

gpt-5.6-terra
Chat

GPT-6 Astra is OpenAI's flagship model for demanding end-to-end work. It is suited for advanced analysis, software engineering, deep research, scientific work, and document creation, with particular strengths in long-horizon

Context

1M

In EGP / 1M

573.02

Out EGP / 1M

2865.08

gpt-6-astra

GPT-6 Astra Pro is the same underlying model as GPT-6 Astra, served with `reasoning.mode` set to `pro` for higher-quality responses on complex tasks. Learn more in OpenAI's docs: https://developers.openai.com/api/docs/guides/reasoning#reasoning-mode

Context

1M

In EGP / 1M

573.02

Out EGP / 1M

2865.08

gpt-6-astra-pro
Chat

GPT-6 Luna is the fast, cost-efficient model in OpenAI's GPT-6 series, positioned below GPT-6 Sol. It is suited for high-volume and latency-sensitive workloads such as chat, classification, and lightweight agentic

Context

1M

In EGP / 1M

5.73

Out EGP / 1M

28.65

gpt-6-luna

GPT-6 Luna Pro is the same underlying model as GPT-6 Luna, served with `reasoning.mode` set to `pro` for higher-quality responses on complex tasks. Learn more in OpenAI's docs: https://developers.openai.com/api/docs/guides/reasoning#reasoning-mode

Context

1M

In EGP / 1M

5.73

Out EGP / 1M

28.65

gpt-6-luna-pro
Chat

GPT-6 Sol is the cost-efficient high-end model in OpenAI's GPT-6 series, positioned below the flagship GPT-6 Astra and above the fast GPT-6 Luna tier. It is suited for demanding professional

Context

1M

In EGP / 1M

114.60

Out EGP / 1M

573.02

gpt-6-sol
Chat

GPT-6 Sol Pro is the same underlying model as GPT-6 Sol, served with `reasoning.mode` set to `pro` for higher-quality responses on complex tasks. Learn more in OpenAI's docs: https://developers.openai.com/api/docs/guides/reasoning#reasoning-mode

Context

1M

In EGP / 1M

114.60

Out EGP / 1M

573.02

gpt-6-sol-pro
Chat

GPT-6.1 Sol is an upgrade to GPT-6 Sol from OpenAI, positioned below the flagship GPT-6 Astra in the GPT-6 series. It is suited for agentic coding, computer use, document-heavy professional

Context

1M

In EGP / 1M

114.60

Out EGP / 1M

573.02

gpt-6.1-sol

GPT-6.1 Sol Pro is the same underlying model as GPT-6.1 Sol, served with `reasoning.mode` set to `pro` for higher-quality responses on complex tasks. **Cost note:** pro mode spends far more

Context

1M

In EGP / 1M

114.60

Out EGP / 1M

573.02

gpt-6.1-sol-pro

gpt-oss-safeguard-20b is a safety reasoning model from OpenAI built upon gpt-oss-20b. This open-weight, 21B-parameter Mixture-of-Experts (MoE) model offers lower latency for safety tasks like content classification, LLM filtering, and trust

Context

131K

In EGP / 1M

4.30

Out EGP / 1M

17.19

gpt-oss-safeguard-20b
OpenAI
Chat

The latest and strongest model family from OpenAI, o1 is designed to spend more time thinking before responding. The o1 model series is trained with large-scale reinforcement learning to reason

Context

200K

In EGP / 1M

859.52

Out EGP / 1M

3438.09

OpenAI
Chat

The o1 series of models are trained with reinforcement learning to think before they answer and perform complex reasoning. The o1-pro model uses more compute to think harder and provide

Context

200K

In EGP / 1M

8595.23

Out EGP / 1M

34380.92

o1-pro
OpenAI
Chat

o3 is a well-rounded and powerful model across domains. It sets a new standard for math, science, coding, and visual reasoning tasks. It also excels at technical writing and instruction-following

Context

200K

In EGP / 1M

114.60

Out EGP / 1M

458.41

OpenAI
Chat

OpenAI o3-mini is a cost-efficient language model optimized for STEM reasoning tasks, particularly excelling in science, mathematics, and coding. This model supports the `reasoning_effort` parameter, which can be set to

Context

200K

In EGP / 1M

63.03

Out EGP / 1M

252.13

o3-mini
Chat

OpenAI o3-mini-high is the same model as o3-mini with reasoning_effort set to high. o3-mini is a cost-efficient language model optimized for STEM reasoning tasks, particularly excelling in science, mathematics, and

Context

200K

In EGP / 1M

63.03

Out EGP / 1M

252.13

o3-mini-high
OpenAI
Chat

The o-series of models are trained with reinforcement learning to think before they answer and perform complex reasoning. The o3-pro model uses more compute to think harder and provide consistently

Context

200K

In EGP / 1M

1146.03

Out EGP / 1M

4584.12

o3-pro
OpenAI
Chat

OpenAI o4-mini is a compact reasoning model in the o-series, optimized for fast, cost-efficient performance while retaining strong multimodal and agentic capabilities. It supports tool use and demonstrates competitive reasoning

Context

200K

In EGP / 1M

63.03

Out EGP / 1M

252.13

o4-mini
Chat

OpenAI o4-mini-high is the same model as o4-mini with reasoning_effort set to high. OpenAI o4-mini is a compact reasoning model in the o-series, optimized for fast, cost-efficient performance while retaining

Context

200K

In EGP / 1M

63.03

Out EGP / 1M

252.13

o4-mini-high
PrunaAI
imageProprietary

P-Image is a state-of-the-art real-time generation model with exceptional text rendering, fine-detail accuracy, and rock-solid prompt adherence. It’s built for instant creativity at high-fidelity images in about one second at a fraction of typical model costs.

Context

—

Price

0.2865 EGP per image

EmbeddingOpen weights

We present a sentence similarity model based on the Sentence Transformers architecture, which maps sentences to a 384-dimensional dense vector space. The model uses a pre-trained BERT encoder and applies mean pooling on top of the contextualized word embeddings to obtain sentence embeddings. We evaluate the model on the Sentence Embeddings Benchmark.

Context

512

In EGP / 1M

0.29

Out EGP / 1M

—

Pareto

Live
unbiased
Chat

Pareto is a multimodal composite model built for research, coding, and agentic workflows, while delivering frontier-level performance across a broad range of general-purpose tasks.

Context

262K

In EGP / 1M

143.25

Out EGP / 1M

429.76

pareto
Chat

Perceptron Mk1 (Mark One) is Perceptron's highest-quality vision-language model for video and embodied reasoning.** It accepts image and video inputs paired with natural language queries, and produces detailed visual understanding

Context

33K

In EGP / 1M

8.60

Out EGP / 1M

85.95

perceptron-mk1
Chat

Perceptron Mk1.5 is Perceptron's embodied reasoning model for physical agents. It accepts text, image, video, and audio input, and answers with text plus optional structured annotations: points, boxes, polygons, tracks,

Context

37K

In EGP / 1M

8.60

Out EGP / 1M

85.95

perceptron-mk1.5
perplexity
Chat

Sonar is lightweight, affordable, fast, and simple to use — now featuring citations and the ability to customize sources. It is designed for companies seeking to integrate lightweight question-and-answer features

Context

127K

In EGP / 1M

57.30

Out EGP / 1M

57.30

sonar
perplexity
Chat

Note: Sonar Pro pricing includes Perplexity search pricing. See details here For enterprises seeking more advanced capabilities, the Sonar Pro API can handle in-depth, multi-step queries with added extensibility, like

Context

200K

In EGP / 1M

171.90

Out EGP / 1M

859.52

sonar-pro
Chat

Exclusively available on the OpenRouter API, Sonar Pro's new Pro Search mode is Perplexity's most advanced agentic search system. It is designed for deeper reasoning and analysis. Pricing is based

Context

200K

In EGP / 1M

171.90

Out EGP / 1M

859.52

sonar-pro-search

Note: Sonar Pro pricing includes Perplexity search pricing. See details here Sonar Reasoning Pro is a premier reasoning model powered by DeepSeek R1 with Chain of Thought (CoT). Designed for

Context

128K

In EGP / 1M

114.60

Out EGP / 1M

458.41

sonar-reasoning-pro

phi-4

Live
Microsoft
ChatOpen weights

Phi-4 is a model built upon a blend of synthetic datasets, data from filtered public domain websites, and acquired academic books and Q&A datasets. The goal of this approach was to ensure that small capable models were trained with data focused on high quality and advanced reasoning.

Context

16K

In EGP / 1M

4.01

Out EGP / 1M

8.02

Chat

Laguna S 2.1 is the latest coding agent model from Poolside. Laguna S 2.1 is a 118B total parameter model with 8B active parameters, scoring 70.2% on Terminal-Bench 2.1 and

Context

1M

In EGP / 1M

5.16

Out EGP / 1M

10.31

laguna-s-2.1
Chat

Laguna XS 2.1 is the latest coding agent model in the 33B-A3B category from Poolside and a step forward from their Laguna XS.2 model (released in April 2026). It combines

Context

262K

In EGP / 1M

3.44

Out EGP / 1M

6.88

laguna-xs-2.1

Bonsai 2 27B is a 27B-parameter reasoning model from PrismML derived from Qwen3.8-27B. It supports coding, mathematics, tool calling, and image understanding with a 262K-token context window. Ternary compression shrinks

Context

262K

In EGP / 1M

4.30

Out EGP / 1M

28.65

ternary-bonsai-2-27b
Alibaba
imageProprietary

Compared with the Plus series, it significantly reduces the “AI-like” feel in generated images, enhancing their realism. It delivers more lifelike material textures for human subjects, finer and more detailed natural textures, and more visually appealing text rendering.

Context

—

Price

4.30 EGP per image

Chat

Qwen Plus 0728, based on the Qwen3 foundation model, is a 1 million context hybrid reasoning model with a balanced performance, speed, and cost combination.

Context

1M

In EGP / 1M

14.90

Out EGP / 1M

44.70

qwen-plus-2025-07-28

Qwen2.5-VL is proficient in recognizing common objects such as flowers, birds, fish, and insects. It is also highly capable of analyzing texts, charts, icons, graphics, and layouts within images.

Context

128K

In EGP / 1M

45.84

Out EGP / 1M

57.30

qwen2.5-vl-72b-instruct
Chat

Qwen3-235B-A22B is a 235B parameter mixture-of-experts (MoE) model developed by Qwen, activating 22B parameters per forward pass. It supports seamless switching between a "thinking" mode for complex reasoning, math, and

Context

131K

In EGP / 1M

26.07

Out EGP / 1M

104.29

qwen3-235b-a22b

Qwen3-235B-A22B-Instruct-2507 is a multilingual, instruction-tuned mixture-of-experts language model based on the Qwen3-235B architecture, with 22B active parameters per forward pass. It is optimized for general-purpose text generation, including instruction following,

Context

262K

In EGP / 1M

5.01

Out EGP / 1M

20.06

qwen3-235b-a22b-2507

Qwen3-235B-A22B-Thinking-2507 is a high-performance, open-weight Mixture-of-Experts (MoE) language model optimized for complex reasoning tasks. It activates 22B of its 235B parameters per forward pass and natively supports up to 262,144

Context

131K

In EGP / 1M

13.18

Out EGP / 1M

131.79

qwen3-235b-a22b-thinking-2507

Qwen3-30B-A3B-Instruct-2507 is a 30.5B-parameter mixture-of-experts language model from Qwen, with 3.3B active parameters per inference. It operates in non-thinking mode and is designed for high-quality instruction following, multilingual understanding, and

Context

128K

In EGP / 1M

2.76

Out EGP / 1M

11.06

qwen3-30b-a3b-instruct-2507

Qwen3-30B-A3B-Thinking-2507 is a 30B parameter Mixture-of-Experts reasoning model optimized for complex tasks requiring extended multi-step thinking. The model is designed specifically for “thinking mode,” where internal reasoning traces are separated

Context

82K

In EGP / 1M

11.46

Out EGP / 1M

137.52

qwen3-30b-a3b-thinking-2507
Alibaba
Chat

Qwen3-8B is a dense 8.2B parameter causal language model from the Qwen3 series, designed for both reasoning-heavy tasks and efficient dialogue. It supports seamless switching between "thinking" mode for math,

Context

131K

In EGP / 1M

6.70

Out EGP / 1M

26.07

qwen3-8b

Qwen3-Coder-30B-A3B-Instruct is a 30.5B parameter Mixture-of-Experts (MoE) model with 128 experts (8 active per forward pass), designed for advanced code generation, repository-scale understanding, and agentic tool use. Built on the

Context

262K

In EGP / 1M

4.01

Out EGP / 1M

16.04

qwen3-coder-30b-a3b-instruct

Qwen3-Coder-480B-A35B-Instruct is a Mixture-of-Experts (MoE) code generation model developed by the Qwen team. It is optimized for agentic coding tasks such as function calling, tool use, and long-context reasoning over

Context

262K

In EGP / 1M

17.19

Out EGP / 1M

57.30

qwen3-coder

Qwen3 Coder Flash is Alibaba's fast and cost efficient version of their proprietary Qwen3 Coder Plus. It is a powerful coding agent model specializing in autonomous programming via tool calling

Context

1M

In EGP / 1M

11.17

Out EGP / 1M

55.87

qwen3-coder-flash
Chat

Qwen3-Coder-Next is an open-weight causal language model optimized for coding agents and local development workflows. It uses a sparse MoE design with 80B total parameters and only 3B activated per

Context

262K

In EGP / 1M

6.88

Out EGP / 1M

45.84

qwen3-coder-next
Chat

Qwen3 Coder Plus is Alibaba's proprietary version of the Open Source Qwen3 Coder 480B A35B. It is a powerful coding agent model specializing in autonomous programming via tool calling and

Context

1M

In EGP / 1M

37.25

Out EGP / 1M

186.23

qwen3-coder-plus
ChatProprietary

Qwen3-Max-Thinking is the flagship reasoning model in the Qwen3 series, designed for high-stakes cognitive tasks that require deep, multi-step reasoning. By significantly scaling model capacity and reinforcement learning compute, it

Context

256K

In EGP / 1M

68.76

Out EGP / 1M

343.81

qwen3-max-thinking

Qwen3-Next-80B-A3B-Thinking is a reasoning-first chat model in the Qwen3-Next line that outputs structured “thinking” traces by default. It’s designed for hard multi-step problems; math proofs, code synthesis/debugging, logic, and agentic

Context

262K

In EGP / 1M

8.60

Out EGP / 1M

68.76

qwen3-next-80b-a3b-thinking
ChatOpen weights

Qwen3-VL-30B-A3B-Instruct is a multimodal model that unifies strong text generation with visual understanding for images and videos. Its Instruct variant optimizes instruction-following for general multimodal tasks. It excels in perception

Context

262K

In EGP / 1M

8.60

Out EGP / 1M

34.38

qwen3-vl-30b-a3b-instruct

Qwen3-VL-30B-A3B-Thinking is a multimodal model that unifies strong text generation with visual understanding for images and videos. Its Thinking variant enhances reasoning in STEM, math, and complex tasks. It excels

Context

131K

In EGP / 1M

11.46

Out EGP / 1M

137.52

qwen3-vl-30b-a3b-thinking

Qwen3-VL-32B-Instruct is a large-scale multimodal vision-language model designed for high-precision understanding and reasoning across text, images, and video. With 32 billion parameters, it combines deep visual perception with advanced text

Context

131K

In EGP / 1M

5.96

Out EGP / 1M

23.84

qwen3-vl-32b-instruct

Qwen3-VL-8B-Instruct is a multimodal vision-language model from the Qwen3-VL series, built for high-fidelity understanding and reasoning across text, images, and video. It features improved multimodal fusion with Interleaved-MRoPE for long-horizon

Context

131K

In EGP / 1M

6.70

Out EGP / 1M

26.07

qwen3-vl-8b-instruct

Qwen3-VL-8B-Thinking is the reasoning-optimized variant of the Qwen3-VL-8B multimodal model, designed for advanced visual and textual reasoning across complex scenes, documents, and temporal sequences. It integrates enhanced multimodal alignment and

Context

131K

In EGP / 1M

10.31

Out EGP / 1M

120.33

qwen3-vl-8b-thinking
ChatOpen weights

The Qwen3.5 series 397B-A17B native vision-language model is built on a hybrid architecture that integrates a linear attention mechanism with a sparse mixture-of-experts model, achieving higher inference efficiency. It delivers

Context

262K

In EGP / 1M

25.79

Out EGP / 1M

171.90

qwen3.5-397b-a17b

The Qwen3.5 native vision-language series Plus models are built on a hybrid architecture that integrates linear attention mechanisms with sparse mixture-of-experts models, achieving higher inference efficiency. In a variety of

Context

1M

In EGP / 1M

14.90

Out EGP / 1M

89.39

qwen3.5-plus-02-15

Qwen3.5 Plus (April 2026) is a large-scale multimodal language model from Alibaba. It accepts text, image, and video input and produces text output, with a 1M token context window. This

Context

1M

In EGP / 1M

17.19

Out EGP / 1M

103.14

qwen3.5-plus-20260420

The Qwen3.5 122B-A10B native vision-language model is built on a hybrid architecture that integrates a linear attention mechanism with a sparse mixture-of-experts model, achieving higher inference efficiency. In terms of

Context

262K

In EGP / 1M

14.90

Out EGP / 1M

119.19

qwen3.5-122b-a10b
ChatOpen weights

The Qwen3.5 27B native vision-language Dense model incorporates a linear attention mechanism, delivering fast response times while balancing inference speed and performance. Its overall capabilities are comparable to those of

Context

262K

In EGP / 1M

14.90

Out EGP / 1M

148.98

qwen3.5-27b
ChatOpen weights

The Qwen3.5 Series 35B-A3B is a native vision-language model designed with a hybrid architecture that integrates linear attention mechanisms and a sparse mixture-of-experts model, achieving higher inference efficiency. Its overall

Context

262K

In EGP / 1M

8.02

Out EGP / 1M

57.30

qwen3.5-35b-a3b
Alibaba
ChatOpen weights

Qwen3.5-9B is a multimodal foundation model from the Qwen3.5 family, designed to deliver strong reasoning, coding, and visual understanding in an efficient 9B-parameter architecture. It uses a unified vision-language design

Context

262K

In EGP / 1M

5.73

Out EGP / 1M

8.60

qwen3.5-9b
Chat

The Qwen3.5 native vision-language Flash models are built on a hybrid architecture that integrates a linear attention mechanism with a sparse mixture-of-experts model, achieving higher inference efficiency. Compared to the

Context

1M

In EGP / 1M

3.72

Out EGP / 1M

14.90

qwen3.5-flash-02-23
ChatOpen weights

Qwen3.6 27B is a dense 27-billion-parameter language model from the Qwen Team at Alibaba, released in April 2026. It features hybrid multimodal capabilities — accepting text, image, and video inputs

Context

262K

In EGP / 1M

18.34

Out EGP / 1M

183.36

qwen3.6-27b
ChatOpen weights

Qwen3.6-35B-A3B is an open-weight multimodal model from Alibaba Cloud with 35 billion total parameters and 3 billion active parameters per token. It uses a hybrid sparse mixture-of-experts architecture combining Gated

Context

262K

In EGP / 1M

5.73

Out EGP / 1M

54.44

qwen3.6-35b-a3b
Chat

Qwen3.6 Flash is a fast, efficient language model from Alibaba's Qwen 3.6 series. It supports text, image, and video input with a 1M token context window. Tiered pricing kicks in

Context

1M

In EGP / 1M

10.74

Out EGP / 1M

64.46

qwen3.6-flash

Qwen3.6-Max-Preview is a proprietary frontier model from Alibaba Cloud built on a sparse mixture-of-experts architecture with approximately 1 trillion total parameters. It is optimized for agentic coding, tool use, and

Context

262K

In EGP / 1M

58.85

Out EGP / 1M

353.09

qwen3.6-max-preview
Chat

Qwen 3.6 Plus builds on a hybrid architecture that combines efficient linear attention with sparse mixture-of-experts routing, enabling strong scalability and high-performance inference. Compared to the 3.5 series, it delivers

Context

1M

In EGP / 1M

18.62

Out EGP / 1M

111.74

qwen3.6-plus
Chat

Qwen3.7 Flash is a vision-language reasoning model from Alibaba. It is suited for multimodal agents, visual coding, search, and computer interaction, with strengths in object recognition, spatial understanding, and real-world

Context

1M

In EGP / 1M

1.72

Out EGP / 1M

7.45

qwen3.7-flash
ChatProprietary

Qwen3.7-Max is the flagship model in Alibaba's Qwen3.7 series. It supports text input and output and is designed for agent-centric workloads, with particular strengths in coding, office and productivity tasks,

Context

256K

In EGP / 1M

143.25

Out EGP / 1M

429.76

qwen3.7-max
Chat

Qwen3.7-Plus is a cost-effective model in Alibaba's Qwen3.7 series. It supports text and image input with text output, building on the series' text capabilities with a comprehensive upgrade to its

Context

1M

In EGP / 1M

18.34

Out EGP / 1M

73.35

qwen3.7-plus
ChatOpen weights

Qwen3.8 2.4T A95B is an open-weight sparse mixture-of-experts model from Qwen and the open-weight variant of Qwen3.8 Max, with 95 billion active parameters out of 2.4 trillion total. It is

Context

262K

In EGP / 1M

114.60

Out EGP / 1M

343.81

qwen3.8-2.4t-a95b
ChatOpen weights

Qwen3.8 27B is an open-weight dense vision-language model from Qwen. It is suited for coding, professional workflows, research, multimodal interaction, and long-running agent tasks, with flexible thinking that can be

Context

262K

In EGP / 1M

11.46

Out EGP / 1M

143.25

qwen3.8-27b
ChatProprietary

Qwen3.8 Flash is a multimodal reasoning model from Alibaba. It is suited for coding assistance, agentic workflows, visual understanding, document and codebase analysis, desktop interaction, chart analysis, and long-video analysis.

Context

1M

In EGP / 1M

6.48

Out EGP / 1M

21.89

qwen3.8-flash

Qwen3.8 Max 0902 is an updated snapshot of Qwen3.8 Max from Alibaba's Qwen team. It is a 2.4-trillion-parameter mixture-of-experts model that accepts text, image, and video input and returns text,

Context

1M

In EGP / 1M

114.60

Out EGP / 1M

343.81

qwen3.8-max-0902

Qwen3.8 Max Prime is a higher-throughput variant of Qwen3.8 Max from Alibaba's Qwen team, served as a separate SKU at a higher price point. It accepts text, image, and video

Context

1M

In EGP / 1M

229.21

Out EGP / 1M

687.62

qwen3.8-max-prime

Qwen3.8 Omni Flash is an omni-modal reasoning model from Alibaba, the first Qwen model built around agentic capabilities with native audio-video understanding. It is suited for audio-video analysis and summarization,

Context

1M

In EGP / 1M

8.60

Out EGP / 1M

26.93

qwen3.8-omni-flash
Chat

Qwen2.5 72B is the latest series of Qwen large language models. Qwen2.5 brings the following improvements upon Qwen2: - Significantly more knowledge and has greatly improved capabilities in coding and

Context

33K

In EGP / 1M

20.63

Out EGP / 1M

22.92

qwen-2.5-72b-instruct

Qwen2.5-Coder is the latest series of Code-Specific Qwen large language models (formerly known as CodeQwen). Qwen2.5-Coder brings the following improvements upon CodeQwen1.5: - Significantly improvements in **code generation**, **code reasoning**

Context

33K

In EGP / 1M

37.82

Out EGP / 1M

57.30

qwen-2.5-coder-32b-instruct
ChatOpen weights

Qwen2.5 is a model pretrained on a large-scale dataset of up to 18 trillion tokens, offering significant improvements in knowledge, coding, mathematics, and instruction following compared to its predecessor Qwen2. The model also features enhanced capabilities in generating long texts, understanding structured data, and generating structured outputs, while supporting multilingual capabilities for over 29 languages.

Context

33K

In EGP / 1M

20.63

Out EGP / 1M

22.92

ChatOpen weights

Qwen2.5 is the latest series of Qwen large language models. For Qwen2.5, we release a number of base language models and instruction-tuned language models ranging from 0.5 to 72 billion parameters. Qwen2.5 brings the following improvements upon Qwen2:

Context

4K

In EGP / 1M

100.00

Out EGP / 1M

500.00

Qwen/Qwen2.5-7B-Instruct
Alibaba
ChatOpen weights

Qwen3 is the latest generation of large language models in Qwen series, offering a comprehensive suite of dense and mixture-of-experts (MoE) models. Built upon extensive training, Qwen3 delivers groundbreaking advancements in reasoning, instruction-following, agent capabilities, and multilingual support.

Context

41K

In EGP / 1M

6.88

Out EGP / 1M

13.75

ChatOpen weights

Qwen3-235B-A22B-Instruct-2507 is the updated version of the Qwen3-235B-A22B non-thinking mode, featuring Significant improvements in general capabilities, including instruction following, logical reasoning, text comprehension, mathematics, science, coding and tool usage.

Context

262K

In EGP / 1M

5.16

Out EGP / 1M

31.52

Alibaba
ChatOpen weights

Qwen3 is the latest generation of large language models in Qwen series, offering a comprehensive suite of dense and mixture-of-experts (MoE) models. Built upon extensive training, Qwen3 delivers groundbreaking advancements in reasoning, instruction-following, agent capabilities, and multilingual support

Context

41K

In EGP / 1M

6.88

Out EGP / 1M

28.65

Alibaba
ChatOpen weights

Qwen3 is the latest generation of large language models in Qwen series, offering a comprehensive suite of dense and mixture-of-experts (MoE) models. Built upon extensive training, Qwen3 delivers groundbreaking advancements in reasoning, instruction-following, agent capabilities, and multilingual support

Context

41K

In EGP / 1M

4.58

Out EGP / 1M

16.04

ChatOpen weights

Qwen3-Coder-480B-A35B-Instruct is the Qwen3's most agentic code model, featuring Significant Performance on Agentic Coding, Agentic Browser-Use and other foundational coding tasks, achieving results comparable to Claude Sonnet.

Context

262K

In EGP / 1M

17.19

Out EGP / 1M

57.30

EmbeddingOpen weights

The Qwen3 Embedding model series is the latest proprietary model of the Qwen family, specifically designed for text embedding and ranking tasks. Building upon the dense foundational models of the Qwen3 series, it provides a comprehensive range of text embeddings and reranking models in various sizes (0.6B, 4B, and 8B).

Context

33K

In EGP / 1M

0.57

Out EGP / 1M

—

EmbeddingOpen weights

The Qwen3 Embedding model series is the latest proprietary model of the Qwen family, specifically designed for text embedding and ranking tasks. Building upon the dense foundational models of the Qwen3 series, it provides a comprehensive range of text embeddings and reranking models in various sizes (0.6B, 4B, and 8B).

Context

33K

In EGP / 1M

1.15

Out EGP / 1M

—

EmbeddingOpen weights

The Qwen3 Embedding model series is the latest proprietary model of the Qwen family, specifically designed for text embedding and ranking tasks. Building upon the dense foundational models of the Qwen3 series, it provides a comprehensive range of text embeddings and reranking models in various sizes (0.6B, 4B, and 8B).

Context

33K

In EGP / 1M

0.57

Out EGP / 1M

—

Alibaba
ChatProprietary

The latest flagship model in the Qwen family. State-of-the-art results across a comprehensive suite of benchmarks — including knowledge, reasoning, coding, instruction following, human preference alignment, agent tasks, and multilingual understanding.

Context

256K

In EGP / 1M

68.76

Out EGP / 1M

343.81

ChatOpen weights

Over the past few months, we have observed increasingly clear trends toward scaling both total parameters and context lengths in the pursuit of more powerful and agentic artificial intelligence (AI). We are excited to share our latest advancements in addressing these demands, centered on improving scaling efficiency through innovative model architecture. We call this next-generation foundation models Qwen3-Next.

Context

262K

In EGP / 1M

5.16

Out EGP / 1M

63.03

ChatOpen weights

Meet Qwen3-VL — the most powerful vision-language model in the Qwen series to date. This generation delivers comprehensive upgrades across the board: superior text understanding & generation, deeper visual perception & reasoning, extended context length, enhanced spatial and video dynamics comprehension, and stronger agent interaction capabilities.

Context

262K

In EGP / 1M

11.46

Out EGP / 1M

50.43

Alibaba
ChatProprietary

Qwen's latest 2.4-trillion-parameter MoE flagship model delivering a comprehensive leap in coding and professional work.

Context

256K

In EGP / 1M

94.55

Out EGP / 1M

283.70

rekaai
Chat

Reka Edge is an extremely efficient 7B multimodal vision-language model that accepts image/video+text inputs and generates text outputs. This model is optimized specifically to deliver industry-leading performance in image understanding,

Context

16K

In EGP / 1M

5.73

Out EGP / 1M

5.73

reka-edge
rekaai
Chat

Reka Flash 3 is a general-purpose, instruction-tuned large language model with 21 billion parameters, developed by Reka. It excels at general chat, coding tasks, instruction-following, and function calling. Featuring a

Context

66K

In EGP / 1M

5.73

Out EGP / 1M

11.46

reka-flash-3
Chat

The relace-search model uses 4-12 `view_file` and `grep` tools in parallel to explore a codebase and return relevant files to the user request. In contrast to RAG, relace-search performs agentic

Context

256K

In EGP / 1M

57.30

Out EGP / 1M

171.90

relace-search
respan
Decision

Span-01 is a behavior scoring model from Respan. It reads a conversation span and returns, for each plain-language behavior you define, the probability that the behavior is present. It is

Context

—

In EGP / 1M

1.15

Out EGP / 1M

Free

span-01
stabilityai
imageOpen weights

The SDXL Turbo model, developed by Stability AI, is an optimized, fast text-to-image generative model. It is a distilled version of SDXL 1.0, leveraging Adversarial Diffusion Distillation (ADD) to generate high-quality images in less steps.

Context

—

Price

0.0115 EGP per 1024p image

ByteDance
ChatProprietary

Optimized specifically for multimodal agent scenarios. It features enhanced agent capabilities, upgraded multimodal comprehension, and more flexible context management.

Context

256K

In EGP / 1M

14.33

Out EGP / 1M

114.60

ByteDance
ChatProprietary

Built for the Agent era, it delivers stable performance in complex reasoning and long-horizon tasks, including multi-step planning, visual-text reasoning, video understanding, and advanced analysis.

Context

256K

In EGP / 1M

28.65

Out EGP / 1M

171.90

ByteDance
imageProprietary

Seedream 4.0 is a SOTA multimodal image creation model built on leading architecture. It breaks through the boundaries of traditional text-to-image models by natively supporting text, single-image, and multi-image inputs. Users can freely combine text and images to achieve diverse creative modes within a single model—such as multi-image blending, image editing, and sequentially batch image generation, featuring subject consistency, making image creation more free and controllable.

Context

—

Price

2.29 EGP per image

ByteDance
imageProprietary

The latest image model, delivering better editing consistency, improved multi-image fusion, finer detail control, natural small text and faces, and harmonious, aesthetic visuals.

Context

—

Price

2.29 EGP per image

ByteDance
imageProprietary

ByteDance's flagship image model, with precise interactive editing, stronger multi-reference fusion and notably accurate text rendering at up to 2K.

Context

—

Price

5.67 EGP per image

Chat

Grok 4.20 is a reasoning model from SpaceXAI with industry-leading speed and agentic tool calling capabilities. It combines the lowest hallucination rate on the market with strict prompt adherance, delivering

Context

2M

In EGP / 1M

71.63

Out EGP / 1M

143.25

grok-4.20
Chat

Grok 4.3 is a reasoning model from SpaceXAI. It accepts text and image inputs with text output, and is suited for agentic workflows, instruction-following tasks, and applications requiring high factual

Context

1M

In EGP / 1M

71.63

Out EGP / 1M

143.25

grok-4.3
Chat

Grok 4.5 is a model from SpaceXAI with frontier performance on coding, knowledge work, and STEM.

Context

500K

In EGP / 1M

114.60

Out EGP / 1M

343.81

grok-4.5
Chat

Grok 4.7 is SpaceXAI's flagship model for coding, agentic tasks, and knowledge work, succeeding Grok 4.6. It is particularly strong at long-running software engineering tasks, verifying its own work, and

Context

500K

In EGP / 1M

114.60

Out EGP / 1M

343.81

grok-4.7

Grok Build 0.1 is SpaceXAI’s fast coding model trained specifically for agentic software engineering workflows. It supports text and image inputs with text output, and is optimized for interactive coding

Context

256K

In EGP / 1M

57.30

Out EGP / 1M

114.60

grok-build-0.1
Chat

Step 3.5 Flash is StepFun's most capable open-source foundation model. Built on a sparse Mixture of Experts (MoE) architecture, it selectively activates only 11B of its 196B parameters per token

Context

262K

In EGP / 1M

5.73

Out EGP / 1M

17.19

step-3.5-flash
Chat

Step 3.7 Flash is StepFun's latest high-efficiency multimodal Mixture-of-Experts model. It pairs a 196B-parameter language backbone with a vision encoder for native image and video understanding, activating roughly 11B parameters

Context

256K

In EGP / 1M

11.46

Out EGP / 1M

65.90

step-3.7-flash

Hunyuan-A13B is a 13B active parameter Mixture-of-Experts (MoE) language model developed by Tencent, with a total parameter count of 80B and support for reasoning via Chain-of-Thought. It offers competitive benchmark

Context

131K

In EGP / 1M

8.02

Out EGP / 1M

32.66

hunyuan-a13b-instruct
Chat

Hy-MT2-1.8B is a compact 1.8B-parameter translation model from Tencent. It supports 33 language pairs and five Chinese dialect and minority-language pairs, with workflows for structured, delimiter-based, contextual, glossary-based, and style-guided

Context

8K

In EGP / 1M

2.52

Out EGP / 1M

10.14

hy-mt2-1.8b
Chat

Hy-MT2-30B-A3B is Tencent's flagship translation model in the Hy-MT2 family. It supports 33 language pairs and five Chinese dialect and minority-language pairs, with workflows for structured, delimiter-based, contextual, glossary-based, and

Context

8K

In EGP / 1M

4.24

Out EGP / 1M

16.90

hy-mt2-30b-a3b
tencent
Chat

Hy-MT2-7B is a 7B-parameter translation model from Tencent. It supports 33 language pairs and five Chinese dialect and minority-language pairs, with workflows for structured, delimiter-based, contextual, glossary-based, and style-guided translation.

Context

8K

In EGP / 1M

4.24

Out EGP / 1M

16.90

hy-mt2-7b
tencent
ChatOpen weights

Hy3 is a 295B-parameter Mixture-of-Experts model from Tencent (21B active, 192 experts with top-8 routing) built for reasoning, agentic workflows, and real-world production use. It supports a configurable reasoning effort:

Context

262K

In EGP / 1M

7.45

Out EGP / 1M

30.37

Chat

Hy3 preview is a high-efficiency Mixture-of-Experts model from Tencent designed for agentic workflows and production use. It supports configurable reasoning levels across disabled, low, and high modes, allowing it to

Context

262K

In EGP / 1M

10.31

Out EGP / 1M

34.38

hy3-preview
ChatOpen weights

Tencent: Hy4 preview is a mixture-of-experts model from Tencent, with 49B active parameters out of 770B total. It is designed for coding agents, complex tool-use workflows, and productivity tasks that

Context

1M

In EGP / 1M

47.79

Out EGP / 1M

143.31

hy4-preview
shibing624
EmbeddingOpen weights

A sentence similarity model that can be used for various NLP tasks such as text classification, sentiment analysis, named entity recognition, question answering, and more. It utilizes the CoSENT architecture, which consists of a transformer encoder and a pooling module, to encode input texts into vectors that capture their semantic meaning. The model was trained on the nli_zh dataset and achieved high performance on various benchmark datasets.

Context

512

In EGP / 1M

0.29

Out EGP / 1M

—

thinkingmachines
ChatOpen weights

Inkling is an open-weight multimodal mixture-of-experts model from Thinking Machines Lab, with 41B active parameters out of 975B total. It is designed for general-purpose reasoning, coding, agentic and tool-use systems,

Context

524K

In EGP / 1M

54.44

Out EGP / 1M

232.07

inkling
thinkingmachines
ChatOpen weights

Inkling Small is an open-weight multimodal mixture-of-experts model from Thinking Machines Lab, with 12B active parameters out of 276B total. It is positioned as the smaller, more efficient member of

Context

524K

In EGP / 1M

25.79

Out EGP / 1M

68.76

inkling-small
togethercomputer
Decision

Tev1 4B Experimental is an experimental decision model from Together AI, a supervised fine-tune of Qwen3.5-4B trained to choose one option from a structured state, question, and list of 2-24

Context

33K

In EGP / 1M

2.41

Out EGP / 1M

Free

tev1-4b-experimental
typesafe
Decision

Jev is a structured decision model from TypeSafe, and the first of its System One models. System One models make fast, structured decisions for software, returning a typed choice rather

Context

32K

In EGP / 1M

2.41

Out EGP / 1M

Free

jev-1.13
Decision

Solar Decide is Upstage's structured decision model, served as a System One endpoint on Solar Mini 4. Send a state along with typed questions, and it returns a choice, a

Context

524K

In EGP / 1M

2.87

Out EGP / 1M

Free

solar-decide
Chat

Solar Mini 4 is Upstage's compact, cost-efficient language model, a 35B-parameter mixture-of-experts with 3B active parameters and a 524K context window. It is built for agentic use cases where response

Context

524K

In EGP / 1M

2.87

Out EGP / 1M

11.46

solar-mini4
Chat

Solar Pro 3 is Upstage's powerful Mixture-of-Experts (MoE) language model. With 102B total parameters and 12B active parameters per forward pass, it delivers exceptional performance while maintaining computational efficiency. Optimized

Context

131K

In EGP / 1M

8.60

Out EGP / 1M

34.38

solar-pro-3
Chat

Solar Pro 4 is Upstage's cost-efficient large language model, featuring a 524K context window. It is built for long-horizon tasks and agentic workflows, with strong capabilities in office productivity, document-intensive

Context

524K

In EGP / 1M

5.16

Out EGP / 1M

20.63

solar-pro4
ChatOpen weights

Voxtral Small is an enhancement of Mistral Small 3, incorporating state-of-the-art audio input capabilities while retaining best-in-class text performance. It excels at speech transcription, translation and audio understanding.

Context

33K

In EGP / 1M

5.73

Out EGP / 1M

17.19

Wan-AI
imageProprietary

Wan2.6 text to image, Upgraded visual quality, aesthetics, and instruction-following deliver precise style control, realistic portraits, long-text understanding, and broad historical/cultural IP coverage, enabling high-quality, highly expressive visual generation.

Context

—

Price

1.72 EGP per image

Microsoft
Chat

WizardLM-2 8x22B is Microsoft AI's most advanced Wizard model. It demonstrates highly competitive performance compared to leading proprietary models, and it consistently outperforms all existing state-of-the-art opensource models. It is

Context

66K

In EGP / 1M

35.53

Out EGP / 1M

35.53

wizardlm-2-8x22b
writer
Chat

Palmyra X5 is Writer's most advanced model, purpose-built for building and scaling AI agents across the enterprise. It delivers industry-leading speed and efficiency on context windows up to 1 million

Context

1M

In EGP / 1M

34.38

Out EGP / 1M

343.81

palmyra-x5
~x-ai
Chat

This model always redirects to the latest Grok model from xAI.

Context

500K

In EGP / 1M

114.60

Out EGP / 1M

343.81

grok-latest
xiaomi
Chat

MiMo-V2.5 is a native omnimodal model by Xiaomi. It delivers Pro-level agentic performance at roughly half the inference cost, while surpassing MiMo-V2-Omni in multimodal perception across image and video understanding

Context

1M

In EGP / 1M

8.02

Out EGP / 1M

16.04

mimo-v2.5
Chat

MiMo-V2.5-Pro is Xiaomi’s flagship model, delivering strong performance in general agentic capabilities, complex software engineering, and long-horizon tasks, with top rankings on benchmarks such as ClawEval, GDPVal, and SWE-bench Pro

Context

1M

In EGP / 1M

24.93

Out EGP / 1M

49.85

mimo-v2.5-pro
ChatOpen weights

MiMo-V2.6-Flash is an open-source foundation model developed by Xiaomi. Built on a Mixture-of-Experts architecture with 309B total parameters and 15B activated per token, it employs a hybrid attention mechanism for

Context

1M

In EGP / 1M

8.02

Out EGP / 1M

16.04

mimo-v2.6-flash
ChatOpen weights

MiMo-V2.6-Pro is the flagship foundation model developed by Xiaomi. Built at a scale of over 1T parameters, it is designed to push the ceiling of capability for the most demanding

Context

1M

In EGP / 1M

24.64

Out EGP / 1M

49.85

mimo-v2.6-pro

MiMo-V2.6-Pro-UltraSpeed is the fast speed edition of Xiaomi's flagship foundation model, MiMo-V2.6-Pro. Built from the same 1T MiMo-V2.6-Pro checkpoint, it matches the original model in quality while delivering roughly 10x

Context

1M

In EGP / 1M

249.26

Out EGP / 1M

498.52

mimo-v2.6-pro-ultraspeed
z-ai
Chat

GLM-4.5 is our latest flagship foundation model, purpose-built for agent-based applications. It leverages a Mixture-of-Experts (MoE) architecture and supports a context length of up to 128k tokens. GLM-4.5 delivers significantly

Context

131K

In EGP / 1M

34.38

Out EGP / 1M

126.06

glm-4.5
Chat

GLM-4.5-Air is the lightweight variant of our latest flagship model family, also purpose-built for agent-centric applications. Like GLM-4.5, it adopts the Mixture-of-Experts (MoE) architecture but with a more compact parameter

Context

131K

In EGP / 1M

7.45

Out EGP / 1M

48.71

glm-4.5-air
z-ai
Chat

GLM-4.5V is a vision-language foundation model for multimodal agent applications. Built on a Mixture-of-Experts (MoE) architecture with 106B parameters and 12B activated parameters, it achieves state-of-the-art results in video understanding,

Context

66K

In EGP / 1M

34.38

Out EGP / 1M

103.14

glm-4.5v
z-ai
Chat

GLM-4.6V is a large multimodal model designed for high-fidelity visual understanding and long-context reasoning across images, documents, and mixed media. It supports up to 128K tokens, processes complex page layouts

Context

131K

In EGP / 1M

17.19

Out EGP / 1M

51.57

glm-4.6v
z-ai
ChatOpen weights

GLM-4.7 is Z.ai’s latest flagship model, featuring upgrades in two key areas: enhanced programming capabilities and more stable multi-step reasoning/execution. It demonstrates significant improvements in executing complex agent tasks while

Context

203K

In EGP / 1M

22.92

Out EGP / 1M

100.28

glm-4.7
Chat

As a 30B-class SOTA model, GLM-4.7-Flash offers a new option that balances performance and efficiency. It is further optimized for agentic coding use cases, strengthening coding capabilities, long-horizon task planning,

Context

131K

In EGP / 1M

3.47

Out EGP / 1M

22.92

glm-4.7-flash
z-ai
Chat

GLM-5 is Z.ai’s flagship open-source foundation model engineered for complex systems design and long-horizon agent workflows. Built for expert developers, it delivers production-grade performance on large-scale programming tasks, rivaling leading

Context

198K

In EGP / 1M

34.38

Out EGP / 1M

110.02

glm-5
Chat

GLM-5 Turbo is a new model from Z.ai designed for fast inference and strong performance in agent-driven environments such as OpenClaw scenarios. It is deeply optimized for real-world agent workflows

Context

203K

In EGP / 1M

68.76

Out EGP / 1M

229.21

glm-5-turbo
z-ai
Chat

GLM-5.1 delivers a major leap in coding capability, with particularly significant gains in handling long-horizon tasks. Unlike previous models built around minute-level interactions, GLM-5.1 can work independently and continuously on

Context

203K

In EGP / 1M

55.27

Out EGP / 1M

173.72

glm-5.1
z-ai
ChatOpen weights

GLM 5.2 is a large-scale reasoning model from Z.ai. It supports text input and output with a 1M-token context window, and is suited for long-horizon agent workflows, project-level software engineering,

Context

1M

In EGP / 1M

42.98

Out EGP / 1M

137.52

glm-5.2
z-ai
ChatOpen weights

GLM-5.3 is a large-scale reasoning model from Z.ai, built for complex software engineering and long-horizon agent tasks. It supports text input and output with a 1M-token context window, and improves

Context

1M

In EGP / 1M

51.57

Out EGP / 1M

229.21

glm-5.3
ChatOpen weights

GLM-5.3-Flash is a native multimodal model from Z.ai. It is suited for efficient coding and long-horizon agent tasks. Its hybrid sparse and linear attention architecture maintains accurate long-context behavior while

Context

1M

In EGP / 1M

8.60

Out EGP / 1M

28.65

glm-5.3-flash
Chat

GLM-5.3-FlashX is the high-speed variant of Z.ai's GLM-5.3-Flash, a native multimodal model delivering inference speeds of up to 200 tokens/s. Built on the same hybrid sparse and linear attention architecture

Context

1M

In EGP / 1M

21.20

Out EGP / 1M

71.63

glm-5.3-flashx
Chat

GLM-5.3-Prime is the high-speed variant of Z.ai's GLM-5.3, inheriting its full capabilities while delivering 1.5–2× the output throughput through inference acceleration. It supports text input and output with a 1M-token

Context

1M

In EGP / 1M

160.44

Out EGP / 1M

504.25

glm-5.3-prime
Chat

GLM-5V-Turbo is Z.ai’s first native multimodal agent foundation model, built for vision-based coding and agent-driven tasks. It natively handles image, video, and text inputs, excels at long-horizon planning, complex coding,

Context

203K

In EGP / 1M

68.76

Out EGP / 1M

229.21

glm-5v-turbo
Chat

This model always redirects to the latest model in the GLM Flash family.

Context

1M

In EGP / 1M

1.15

Out EGP / 1M

14.18

glm-flash-latest
~z-ai
Chat

This model always redirects to the latest GLM model from Z.ai.

Context

1M

In EGP / 1M

8.82

Out EGP / 1M

27.73

glm-latest