LLM Reference
Microsoft Foundry

Using Falcon 40B on Microsoft Foundry

Implementation guide · Falcon · Technology Innovation Institute (TII)

ProvisionedOpen Source

Microsoft Foundry exposes Falcon 40B through model ID falcon-40b. Use the setup steps, sourced pricing, capabilities, and official provider links below to validate this route before deployment.

Last refreshed 2026-05-19. Next refresh: weekly.

Quick Start

  1. 1
    Create an account at Microsoft Foundry and generate an API key.
  2. 2
    Use the Microsoft Foundry SDK or REST API to call falcon-40b — see the documentation for request format.
  3. 3
    You'll be billed $1.54/1M input, $1.77/1M output tokens. See full pricing.

Code Examples

See Microsoft Foundry documentation for integration details.

Pricing on Microsoft Foundry

TypePrice (per 1M)
Input tokens$1.54
Output tokens$1.77

Capabilities

Structured Outputs

About Falcon 40B

Falcon 40B is a leading open-source large language model developed by the Technology Innovation Institute in Abu Dhabi, featuring a causal decoder-only architecture with 40 billion parameters. It stands out with its use of rotary positional embeddings, multi-query attention, and FlashAttention, enhancing its contextual understanding and processing efficiency. Trained on 1 trillion tokens using the enriched RefinedWeb dataset, Falcon 40B excels in various natural language processing tasks, ranging from text generation to language translation and question answering. It supports multiple languages and is open under the Apache 2.0 license, promoting both research and commercial use.

Model Specs

Released2023-11-28
Parameters40B
ArchitectureDecoder Only

More Models on Microsoft Foundry

Provider

Microsoft Foundry
Microsoft Foundry

Microsoft

Redmond, Washington, United States