Using Falcon 40B on Replicate API

Implementation guide · Falcon · Technology Innovation Institute (TII)

ServerlessOpen Source

Replicate API exposes Falcon 40B through model ID falcon-40b. Use the setup steps, sourced pricing, capabilities, and official provider links below to validate this route before deployment.

Last refreshed 2026-05-19. Next refresh: weekly.

Quick Start

  1. 1
    Create an account at Replicate API and generate an API key.
  2. 2
    Use the Replicate API SDK or REST API to call falcon-40b — see the documentation for request format.
  3. 3
    You'll be billed $0.65/1M input, $2.75/1M output tokens. See full pricing.

Code Examples

Install
pip install replicate
API key
REPLICATE_API_TOKEN
Model ID
falcon-40b

Replicate uses "owner/model-name" format (e.g. "meta/meta-llama-3-8b-instruct") for the latest version, or "owner/model-name:version-sha" to pin to a specific version. The REST endpoint splits owner and model-name into the path: /v1/models/{owner}/{model-name}/predictions.

import replicate

# reads REPLICATE_API_TOKEN from env
# falcon-40b format: "owner/model-name" (latest version) or "owner/model-name:version-hash"
output = replicate.run(
    "falcon-40b",
    input={"prompt": "Hello"}
)
# Output is a list or generator depending on the model
print("".join(output))

Pricing on Replicate API

TypePrice (per 1M)
Input tokens$0.65
Output tokens$2.75

Capabilities

Structured Outputs

About Falcon 40B

Falcon 40B is a leading open-source large language model developed by the Technology Innovation Institute in Abu Dhabi, featuring a causal decoder-only architecture with 40 billion parameters. It stands out with its use of rotary positional embeddings, multi-query attention, and FlashAttention, enhancing its contextual understanding and processing efficiency. Trained on 1 trillion tokens using the enriched RefinedWeb dataset, Falcon 40B excels in various natural language processing tasks, ranging from text generation to language translation and question answering. It supports multiple languages and is open under the Apache 2.0 license, promoting both research and commercial use.

Model Specs

Released2023-11-28
Parameters40B
ArchitectureDecoder Only

Provider

Replicate API

Replicate

San Francisco, California, United States