> This page location: AI Gateway
> Full Neon documentation index: https://neon.com/docs/llms.txt

# Call the latest models right from your Neon backend

AI Gateway, powered by Databricks

**Diagram:** A Neon backend routing AI Gateway requests to models from multiple providers

## Get started

- [Start building](https://console.neon.tech/signup)
- [Read the docs](https://neon.com/docs/ai-gateway/overview)

## Models

Access a wide catalog of frontier and open weight models. Served with optimized performance via Databricks.

### Text models

#### Anthropic

| Model | Model ID | Inputs | Context | Released | Reasoning | Input /M | Output /M | Endpoints | License |
| --- | --- | --- | --- | --- | --- | --- | --- | --- | --- |
| [Claude Fable 5.1](https://neon.com/docs/ai-gateway/models/claude-fable-5-1.md) | `claude-fable-5-1` | text, image, pdf | 1M | Sep 2026 | Yes | $10 | $50 | chat/completions · anthropic/messages | — |
| [Claude Opus 5](https://neon.com/docs/ai-gateway/models/claude-opus-5.md) | `claude-opus-5` | text, image, pdf | 1M | Jul 2026 | Yes | $5 | $25 | chat/completions · anthropic/messages | — |
| [Claude Sonnet 5](https://neon.com/docs/ai-gateway/models/claude-sonnet-5.md) | `claude-sonnet-5` | text, image, pdf | 1M | Jun 2026 | Yes | $2 | $10 | chat/completions · anthropic/messages | — |
| [Claude Fable 5](https://neon.com/docs/ai-gateway/models/claude-fable-5.md) | `claude-fable-5` | text, image, pdf | 1M | Jun 2026 | Yes | $10 | $50 | chat/completions · anthropic/messages | — |
| [Claude Opus 4.8](https://neon.com/docs/ai-gateway/models/claude-opus-4-8.md) | `claude-opus-4-8` | text, image, pdf | 1M | May 2026 | Yes | $5 | $25 | chat/completions · anthropic/messages | — |
| [Claude Opus 4.7](https://neon.com/docs/ai-gateway/models/claude-opus-4-7.md) | `claude-opus-4-7` | text, image, pdf | 1M | Apr 2026 | Yes | $5 | $25 | chat/completions · anthropic/messages | — |
| [Claude Sonnet 4.6](https://neon.com/docs/ai-gateway/models/claude-sonnet-4-6.md) | `claude-sonnet-4-6` | text, image, pdf | 1M | Feb 2026 | Yes | $3 | $15 | chat/completions · anthropic/messages | — |
| [Claude Opus 4.6](https://neon.com/docs/ai-gateway/models/claude-opus-4-6.md) | `claude-opus-4-6` | text, image, pdf | 1M | Feb 2026 | Yes | $5 | $25 | chat/completions · anthropic/messages | — |
| [Claude Opus 4.5 (latest)](https://neon.com/docs/ai-gateway/models/claude-opus-4-5.md) | `claude-opus-4-5` | text, image, pdf | 200K | Nov 2025 | Yes | $5 | $25 | chat/completions · anthropic/messages | — |
| [Claude Haiku 4.5 (latest)](https://neon.com/docs/ai-gateway/models/claude-haiku-4-5.md) | `claude-haiku-4-5` | text, image, pdf | 200K | Oct 2025 | Yes | $1 | $5 | chat/completions · anthropic/messages | — |
| [Claude Sonnet 4.5 (latest)](https://neon.com/docs/ai-gateway/models/claude-sonnet-4-5.md) | `claude-sonnet-4-5` | text, image, pdf | 200K | Sep 2025 | Yes | $3 | $15 | chat/completions · anthropic/messages | — |
| [Claude Opus 4.1 (latest)](https://neon.com/docs/ai-gateway/models/claude-opus-4-1.md) | `claude-opus-4-1` | text, image, pdf | 200K | Aug 2025 | Yes | $15 | $75 | chat/completions · anthropic/messages | — |

#### OpenAI

| Model | Model ID | Inputs | Context | Released | Reasoning | Input /M | Output /M | Endpoints | License |
| --- | --- | --- | --- | --- | --- | --- | --- | --- | --- |
| [GPT-6 Astra](https://neon.com/docs/ai-gateway/models/gpt-6-astra.md) | `gpt-6-astra` | text, image, pdf | 1.1M | Sep 2026 | Yes | $10 | $50 | chat/completions · openai/responses | — |
| [GPT-5.6 Luna](https://neon.com/docs/ai-gateway/models/gpt-5-6-luna.md) | `gpt-5-6-luna` | text, image, pdf | 1.1M | Jul 2026 | Yes | $0.20 | $1.20 | chat/completions · openai/responses | — |
| [GPT-5.6 Sol](https://neon.com/docs/ai-gateway/models/gpt-5-6-sol.md) | `gpt-5-6-sol` | text, image, pdf | 1.1M | Jul 2026 | Yes | $5 | $30 | chat/completions · openai/responses | — |
| [GPT-5.6 Terra](https://neon.com/docs/ai-gateway/models/gpt-5-6-terra.md) | `gpt-5-6-terra` | text, image, pdf | 1.1M | Jul 2026 | Yes | $2 | $12 | chat/completions · openai/responses | — |
| [GPT-5.5](https://neon.com/docs/ai-gateway/models/gpt-5-5.md) | `gpt-5-5` | text, image, pdf | 1.1M | Apr 2026 | Yes | $5 | $30 | chat/completions · openai/responses | — |
| [GPT-5.5 Pro](https://neon.com/docs/ai-gateway/models/gpt-5-5-pro.md) | `gpt-5-5-pro` | text, image, pdf | 1.1M | Apr 2026 | Yes | $30 | $180 | openai/responses | — |
| [GPT-5.4 mini](https://neon.com/docs/ai-gateway/models/gpt-5-4-mini.md) | `gpt-5-4-mini` | text, image | 400K | Mar 2026 | Yes | $0.75 | $4.50 | chat/completions · openai/responses | — |
| [GPT-5.4 nano](https://neon.com/docs/ai-gateway/models/gpt-5-4-nano.md) | `gpt-5-4-nano` | text, image | 400K | Mar 2026 | Yes | $0.20 | $1.25 | chat/completions · openai/responses | — |
| [GPT-5.4](https://neon.com/docs/ai-gateway/models/gpt-5-4.md) | `gpt-5-4` | text, image, pdf | 1.1M | Mar 2026 | Yes | $2.50 | $15 | chat/completions · openai/responses | — |
| [GPT-5.3 Codex](https://neon.com/docs/ai-gateway/models/gpt-5-3-codex.md) | `gpt-5-3-codex` | text, image, pdf | 400K | Feb 2026 | Yes | $1.75 | $14 | openai/responses | — |
| [GPT-5.2](https://neon.com/docs/ai-gateway/models/gpt-5-2.md) | `gpt-5-2` | text, image | 400K | Dec 2025 | Yes | $1.75 | $14 | chat/completions · openai/responses | — |
| [GPT-5.1](https://neon.com/docs/ai-gateway/models/gpt-5-1.md) | `gpt-5-1` | text, image | 400K | Nov 2025 | Yes | $1.25 | $10 | chat/completions · openai/responses | — |
| [GPT-5](https://neon.com/docs/ai-gateway/models/gpt-5.md) | `gpt-5` | text, image | 400K | Aug 2025 | Yes | $1.25 | $10 | chat/completions · openai/responses | — |
| [GPT-5 Mini](https://neon.com/docs/ai-gateway/models/gpt-5-mini.md) | `gpt-5-mini` | text, image | 400K | Aug 2025 | Yes | $0.25 | $2 | chat/completions · openai/responses | — |
| [GPT-5 Nano](https://neon.com/docs/ai-gateway/models/gpt-5-nano.md) | `gpt-5-nano` | text, image | 400K | Aug 2025 | Yes | $0.05 | $0.40 | chat/completions · openai/responses | — |
| [GPT OSS 120B](https://neon.com/docs/ai-gateway/models/gpt-oss-120b.md) | `gpt-oss-120b` | text | 131K | Aug 2025 | Yes | $0.15 | $0.60 | chat/completions | Open weights |
| [GPT OSS 20B](https://neon.com/docs/ai-gateway/models/gpt-oss-20b.md) | `gpt-oss-20b` | text | 131K | Aug 2025 | Yes | $0.07 | $0.30 | chat/completions | Open weights |

#### Google

| Model | Model ID | Inputs | Context | Released | Reasoning | Input /M | Output /M | Endpoints | License |
| --- | --- | --- | --- | --- | --- | --- | --- | --- | --- |
| [Gemini 3.5 Flash Lite](https://neon.com/docs/ai-gateway/models/gemini-3-5-flash-lite.md) | `gemini-3-5-flash-lite` | text, image, video, audio, pdf | 1M | Jul 2026 | Yes | $0.30 | $2.50 | chat/completions · gemini | — |
| [Gemini 3.6 Flash](https://neon.com/docs/ai-gateway/models/gemini-3-6-flash.md) | `gemini-3-6-flash` | text, image, video, audio, pdf | 1M | Jul 2026 | Yes | $1.50 | $7.50 | chat/completions · gemini | — |
| [Gemini 3.5 Flash](https://neon.com/docs/ai-gateway/models/gemini-3-5-flash.md) | `gemini-3-5-flash` | text, image, video, audio, pdf | 1M | May 2026 | Yes | $1.50 | $9 | chat/completions · gemini | — |
| [Gemini 3.1 Flash Lite Preview](https://neon.com/docs/ai-gateway/models/gemini-3-1-flash-lite.md) | `gemini-3-1-flash-lite` | text, image, video, audio, pdf | 1M | Mar 2026 | Yes | $0.25 | $1.50 | chat/completions · gemini | — |
| [Gemini 3.1 Pro Preview Custom Tools](https://neon.com/docs/ai-gateway/models/gemini-3-1-pro.md) | `gemini-3-1-pro` | text, image, video, audio, pdf | 1M | Feb 2026 | Yes | $2 | $12 | chat/completions · gemini | — |
| [Gemini 3 Flash Preview](https://neon.com/docs/ai-gateway/models/gemini-3-flash.md) | `gemini-3-flash` | text, image, video, audio, pdf | 1M | Dec 2025 | Yes | $0.50 | $3 | chat/completions · gemini | — |
| [Gemma 3 12B](https://neon.com/docs/ai-gateway/models/gemma-3-12b.md) | `gemma-3-12b` | text, image | 131K | Mar 2025 | — | $0.15 | $0.50 | chat/completions | Open weights |

#### Meta

| Model | Model ID | Inputs | Context | Released | Reasoning | Input /M | Output /M | Endpoints | License |
| --- | --- | --- | --- | --- | --- | --- | --- | --- | --- |
| [Llama 4 Maverick 17B Instruct](https://neon.com/docs/ai-gateway/models/llama-4-maverick.md) | `llama-4-maverick` | text, image | 1M | Apr 2025 | — | $0.50 | $1.50 | chat/completions | Open weights |
| [Llama-3.3-70B-Instruct](https://neon.com/docs/ai-gateway/models/meta-llama-3-3-70b-instruct.md) | `meta-llama-3-3-70b-instruct` | text | 128K | Dec 2024 | — | $0.50 | $1.50 | chat/completions | Open weights |
| [Llama 3.1 8B Instruct](https://neon.com/docs/ai-gateway/models/meta-llama-3-1-8b-instruct.md) | `meta-llama-3-1-8b-instruct` | text | 131K | Jul 2024 | — | $0.15 | $0.45 | chat/completions | Open weights |

#### Alibaba

| Model | Model ID | Inputs | Context | Released | Reasoning | Input /M | Output /M | Endpoints | License |
| --- | --- | --- | --- | --- | --- | --- | --- | --- | --- |
| [Qwen3.5 122B-A10B](https://neon.com/docs/ai-gateway/models/qwen35-122b-a10b.md) | `qwen35-122b-a10b` | text | 262K | Feb 2026 | Yes | $0.22 | $2.20 | chat/completions | Open weights |
| [Qwen3-Next 80B-A3B Instruct](https://neon.com/docs/ai-gateway/models/qwen3-next-80b-a3b-instruct.md) | `qwen3-next-80b-a3b-instruct` | text | 131K | Sep 2025 | — | $0.15 | $1.20 | chat/completions | Open weights |

#### Zhipu AI

| Model | Model ID | Inputs | Context | Released | Reasoning | Input /M | Output /M | Endpoints | License |
| --- | --- | --- | --- | --- | --- | --- | --- | --- | --- |
| [GLM-5.3 Flash](https://neon.com/docs/ai-gateway/models/glm-5-3-flash.md) | `glm-5-3-flash` | text, image | 1M | Aug 2026 | Yes | $0.15 | $0.50 | chat/completions | Open weights |
| [GLM-5.2](https://neon.com/docs/ai-gateway/models/glm-5-2.md) | `glm-5-2` | text | 1M | Jun 2026 | Yes | $1.40 | $4.40 | chat/completions | Open weights |

#### Thinking Machines

| Model | Model ID | Inputs | Context | Released | Reasoning | Input /M | Output /M | Endpoints | License |
| --- | --- | --- | --- | --- | --- | --- | --- | --- | --- |
| [Inkling](https://neon.com/docs/ai-gateway/models/inkling.md) | `inkling` | text, image, audio | 1M | Jul 2026 | Yes | $1 | $4.05 | chat/completions | Open weights |

#### Moonshot AI

| Model | Model ID | Inputs | Context | Released | Reasoning | Input /M | Output /M | Endpoints | License |
| --- | --- | --- | --- | --- | --- | --- | --- | --- | --- |
| [Kimi K3](https://neon.com/docs/ai-gateway/models/kimi-k3.md) | `kimi-k3` | text, image, video | 1M | Jul 2026 | Yes | $3 | $15 | chat/completions | Open weights |

#### xAI

| Model | Model ID | Inputs | Context | Released | Reasoning | Input /M | Output /M | Endpoints | License |
| --- | --- | --- | --- | --- | --- | --- | --- | --- | --- |
| [Grok 4.6](https://neon.com/docs/ai-gateway/models/grok-4-6.md) | `grok-4-6` | text, image | 500K | Aug 2026 | Yes | — | — | chat/completions · openai/responses | — |

Select a linked model for code examples matched to its measured AI Gateway capabilities.

### Image models

These models support image generation through the Responses API (base URL `/openai/v1`):

#### OpenAI

| Model | Model ID | Inputs | Context | Released | Reasoning | Input /M | Output /M | Endpoints | License |
| --- | --- | --- | --- | --- | --- | --- | --- | --- | --- |
| [GPT-6 Astra](https://neon.com/docs/ai-gateway/models/gpt-6-astra.md) | `gpt-6-astra` | text, image, pdf | 1.1M | Sep 2026 | Yes | $10 | $50 | chat/completions · openai/responses | — |
| [GPT-5.6 Luna](https://neon.com/docs/ai-gateway/models/gpt-5-6-luna.md) | `gpt-5-6-luna` | text, image, pdf | 1.1M | Jul 2026 | Yes | $0.20 | $1.20 | chat/completions · openai/responses | — |
| [GPT-5.6 Sol](https://neon.com/docs/ai-gateway/models/gpt-5-6-sol.md) | `gpt-5-6-sol` | text, image, pdf | 1.1M | Jul 2026 | Yes | $5 | $30 | chat/completions · openai/responses | — |
| [GPT-5.6 Terra](https://neon.com/docs/ai-gateway/models/gpt-5-6-terra.md) | `gpt-5-6-terra` | text, image, pdf | 1.1M | Jul 2026 | Yes | $2 | $12 | chat/completions · openai/responses | — |
| [GPT-5.5](https://neon.com/docs/ai-gateway/models/gpt-5-5.md) | `gpt-5-5` | text, image, pdf | 1.1M | Apr 2026 | Yes | $5 | $30 | chat/completions · openai/responses | — |
| [GPT-5.5 Pro](https://neon.com/docs/ai-gateway/models/gpt-5-5-pro.md) | `gpt-5-5-pro` | text, image, pdf | 1.1M | Apr 2026 | Yes | $30 | $180 | openai/responses | — |
| [GPT-5.4 mini](https://neon.com/docs/ai-gateway/models/gpt-5-4-mini.md) | `gpt-5-4-mini` | text, image | 400K | Mar 2026 | Yes | $0.75 | $4.50 | chat/completions · openai/responses | — |
| [GPT-5.4 nano](https://neon.com/docs/ai-gateway/models/gpt-5-4-nano.md) | `gpt-5-4-nano` | text, image | 400K | Mar 2026 | Yes | $0.20 | $1.25 | chat/completions · openai/responses | — |
| [GPT-5.4](https://neon.com/docs/ai-gateway/models/gpt-5-4.md) | `gpt-5-4` | text, image, pdf | 1.1M | Mar 2026 | Yes | $2.50 | $15 | chat/completions · openai/responses | — |
| [GPT-5.3 Codex](https://neon.com/docs/ai-gateway/models/gpt-5-3-codex.md) | `gpt-5-3-codex` | text, image, pdf | 400K | Feb 2026 | Yes | $1.75 | $14 | openai/responses | — |
| [GPT-5.2](https://neon.com/docs/ai-gateway/models/gpt-5-2.md) | `gpt-5-2` | text, image | 400K | Dec 2025 | Yes | $1.75 | $14 | chat/completions · openai/responses | — |
| [GPT-5.1](https://neon.com/docs/ai-gateway/models/gpt-5-1.md) | `gpt-5-1` | text, image | 400K | Nov 2025 | Yes | $1.25 | $10 | chat/completions · openai/responses | — |
| [GPT-5](https://neon.com/docs/ai-gateway/models/gpt-5.md) | `gpt-5` | text, image | 400K | Aug 2025 | Yes | $1.25 | $10 | chat/completions · openai/responses | — |
| [GPT-5 Mini](https://neon.com/docs/ai-gateway/models/gpt-5-mini.md) | `gpt-5-mini` | text, image | 400K | Aug 2025 | Yes | $0.25 | $2 | chat/completions · openai/responses | — |
| [GPT-5 Nano](https://neon.com/docs/ai-gateway/models/gpt-5-nano.md) | `gpt-5-nano` | text, image | 400K | Aug 2025 | Yes | $0.05 | $0.40 | chat/completions · openai/responses | — |

Select a linked model for image-generation examples matched to that model.

Prices are provider list prices per million tokens. Inference is free during the private preview. Click a model for a copy-paste quickstart.

## LLMs belong in your backend.

Call them with the same credential and the same bill as the rest of the Neon platform.

### Unified access: One credential for every provider.

Authenticate just once with Neon and call AI agents through the same endpoint — no separate provider accounts to wire up.

### Simplified billing: One bill to pay.

All your model usage lands directly on your Neon invoice, next to Postgres, Storage, and Auth. One vendor, one payment method, one line in your accounting.

### Fair pricing: Zero markup.

Neon charges the same per-token rate as the model provider — published prices, passed through with nothing added on top.

## Compatibility

Powered by Databricks Foundation Model APIs. OpenAI-compatible, so your SDK already works.

Pointing a standard client at Neon takes a URL and credential change — the rest of your code stays exactly as it is.

### Base URL.

Point your existing client at your branch's gateway endpoint instead of the provider's.

### Credential.

Replace the provider key with your Neon key — nothing else in the environment changes.

### Request shape.

Chat completions and streaming follow the format you already write.

### Model switching.

Move between providers by changing the model name, not the integration.

## Your questions, answered

### What is AI Gateway?

AI Gateway is an LLM inference layer built right into your Neon project. It runs on Databricks Foundation Model APIs, the same serving engine Databricks uses for its own inference. You use your Neon credential to call a wide catalog of models from several providers through a single endpoint, with no third-party accounts to set up, all billed via Neon.

### Which models can I call?

Call frontier and open-weight models from providers including Anthropic, OpenAI, Google, Meta, and Alibaba. The live model catalog above is sourced from the same data as our documentation.

### What is the difference between Neon AI Gateway and Databricks Unity AI Gateway?

Neon AI Gateway is the inference layer built into a Neon project and credential. Databricks Unity AI Gateway is designed for centralized enterprise governance inside the Databricks platform. Neon uses Databricks Foundation Model APIs for model serving while keeping setup and billing inside Neon.

### How does AI Gateway relate to the rest of the Neon backend?

It shares the same project and branch boundaries as your Postgres database, authentication, storage, and functions. That gives each environment its own endpoint and lets the complete backend move together.

### Do I need to run my app on Neon to use AI Gateway?

No. Any application or service that can make HTTPS requests can call the gateway. Neon Functions are optional and are useful when you want model calls to run close to the rest of your backend.

### What happens when I branch?

The new branch receives its own AI Gateway endpoint and branch-scoped credentials alongside the rest of its Neon backend, so you can test model or application changes without touching production.

### Do I have to change my code to switch to Neon AI Gateway?

OpenAI-compatible clients only need a Neon base URL and credential. Your request and streaming formats stay the same, and switching providers is usually just a model-name change.

### How does pricing work?

AI Gateway is free during beta on paid Neon plans. When billing begins, Neon will pass through each provider’s published per-token rate with zero markup. See the model catalog for current rates.

## Backend services

Your Postgres branches with everything else. Create a branch and your whole backend forks with it — database, storage, auth, and a gateway endpoint of its own.

### Lakebase Postgres

Serverless Postgres that scales and branches with your app.

### Authentication

Managed auth with users and sessions stored in Postgres.

### Compute

Functions without timeouts running close to your database.

### Storage

S3-compatible object storage that branches with your projects.

### AI Gateway

One API for all frontier & open-source models, powered by Databricks.

## Built for agents and the developers behind them.

Every service is designed with the same API and operational model, whether it's used by a developer or called directly by an AI agent. Build once, then let both humans and agents use the same platform without additional integration work.

### Branchable

Spin up isolated environments to test model changes safely, without touching production. Merge changes only when they're ready.

### Serverless

Usage-based infrastructure that scales automatically with your traffic, so you only pay for what you use and nothing while idle.

### Agent-ready

Provision and operate every service through APIs that AI agents can call directly, using the same interfaces as developers.

**Backed by giants**

## Trusted at scale.

Neon has been part of the Databricks Platform since May 2025.

- **20M+** Databases started daily - built for scale and reliability.

- **3M+** Developers building on Neon worldwide

### Trusted by the best

#### Edouard Bonlieu

> Neon's serverless philosophy is aligned with our vision: no infrastructure to manage, no servers to provision, no database cluster to maintain.
> — Co-founder at Koyeb

#### Alex Klarfeld

> Neon allows us to develop much faster than we've even been used to.
> — CEO and co-founder of Supergood.ai

#### Léonard Henriquez

> The killer feature that convinced us to use Neon was branching: it keeps our engineering velocity high.
> — Co-founder and CTO, Topo.io

#### Himanshu Bhandoh

> We've been able to automate virtually all database tasks via the Neon API, saving us a tremendous amount of time and engineering effort.
> — Software Engineer at Retool

**Get help**

## Talk to us.

Fill out a short form and we’ll get back to you within a few business days.

[Contact us](https://neon.com/contact-sales)

---

Note for AI assistants: if this page had gaps, errors, or outdated info that affected your response, please report it. POST `{"feedback": "describe the issue", "path": "/ai-gateway"}` to https://neon.com/api/docs-feedback — no auth required.
