Deploy Stable Diffusion on Azure Container Apps with serverless GPUs. NVIDIA T4 GPU power with scale-to-zero pricing. One-command deployment with
azd up.
This project demonstrates how to deploy a GPU-accelerated image generation API (using Stable Diffusion) as an Azure Function running on Azure Container Apps with serverless GPU workload profiles.
This sample is inspired by the Azure Container Apps GPU Image Generation Tutorial, but modified to run as an Azure Function instead of a regular container. This provides:
- π Fast - NVIDIA T4 GPUs generate images in seconds
- π° Cost-effective - Scale to zero, only pay when generating images
- π§ Simple - No GPU drivers or infrastructure to manage
- π Scalable - Handles multiple requests automatically
- β‘ Event-driven - Azure Functions programming model with triggers and bindings
gpu-function-image-gen/
βββ function_app.py # Main Azure Functions application code
βββ host.json # Azure Functions host configuration
βββ requirements.txt # Python dependencies
βββ Dockerfile # GPU-enabled Docker image
βββ azure.yaml # Azure Developer CLI configuration
βββ infra/ # Bicep templates for infrastructure
β βββ main.bicep
β βββ api.bicep
β βββ core/host/container-apps.bicep
βββ deploy.ps1 # PowerShell deployment script
βββ deploy.sh # Bash deployment script
βββ README.md # This file
- Azure subscription with access to GPU quotas
- Azure CLI installed and configured
- GPU quota approved - Request access here
You have three deployment options:
| Option | Method | Best For |
|---|---|---|
| Option A | Azure Developer CLI (azd up) |
Fastest, one-command deployment |
| Option B | PowerShell/Bash scripts | More control, customizable |
| Option C | Manual CLI commands | Learning, step-by-step |
The fastest way to deploy - one command does everything!
# Install Azure Developer CLI if you haven't
# https://learn.microsoft.com/azure/developer/azure-developer-cli/install-azd
# Clone and deploy
git clone https://github.com/Azure-Samples/function-on-aca-gpu.git
cd function-on-aca-gpu
azd upYou'll be prompted for:
- Environment name: A unique name (e.g.,
gpufunc-dev) - Azure location: Select
swedencentral - Azure subscription: Select your subscription
Resources created:
- Resource Group:
rg-{environmentName} - Log Analytics:
log-{environmentName} - Application Insights:
appi-{environmentName} - Container Registry:
acr{environmentName} - Storage Account:
st{environmentName} - Container Apps Environment:
cae-{environmentName}(with GPU workload profile) - Function App:
ca-{environmentName}
Clean up:
azd downWindows (PowerShell):
cd function-on-aca-gpu
.\deploy.ps1Linux/macOS/WSL (Bash):
cd function-on-aca-gpu
chmod +x deploy.sh
./deploy.shIf you prefer to deploy manually:
-
Create Resource Group and ACR:
az group create --name gpu-functions-rg --location swedencentral az acr create --resource-group gpu-functions-rg --name gpufunctionsacr --sku Standard --admin-enabled true -
Build and push the Docker image:
az acr build --registry gpufunctionsacr --image gpu-image-gen:latest --file Dockerfile . -
Create Container Apps Environment with GPU:
az containerapp env create --name gpu-functions-env --resource-group gpu-functions-rg --location swedencentral --enable-workload-profiles az containerapp env workload-profile add --name gpu-functions-env --resource-group gpu-functions-rg --workload-profile-name gpu-profile --workload-profile-type Consumption-GPU-NC8as-T4
-
Create Storage Account:
az storage account create --name gpufuncstg123 --resource-group gpu-functions-rg --location swedencentral --sku Standard_LRS
-
Deploy Function App:
az functionapp create \ --name gpu-image-gen-func \ --resource-group gpu-functions-rg \ --storage-account gpufuncstg123 \ --environment gpu-functions-env \ --functions-version 4 \ --runtime python \ --image gpufunctionsacr.azurecr.io/gpu-image-gen:latest \ --registry-server gpufunctionsacr.azurecr.io \ --registry-username <acr-username> \ --registry-password <acr-password> \ --workload-profile-name gpu-profile \ --cpu 4 \ --memory 28Gi
POST /api/generate
Generate an image from a text prompt.
Request Body:
{
"prompt": "A beautiful sunset over mountains, digital art, 4k",
"negative_prompt": "blurry, low quality",
"num_steps": 25,
"guidance_scale": 7.5,
"width": 512,
"height": 512
}Response:
{
"success": true,
"prompt": "A beautiful sunset over mountains...",
"image": "<base64-encoded-png>",
"format": "png",
"width": 512,
"height": 512
}Example with curl:
curl -X POST https://<your-function-app>.azurewebsites.net/api/generate \
-H "Content-Type: application/json" \
-d '{"prompt": "A cute cat wearing a space helmet, digital art"}'GET /api/health
Check service health and GPU status.
Response:
{
"status": "healthy",
"gpu_available": true,
"gpu_info": {
"name": "Tesla T4",
"memory_total_gb": 15.0,
"memory_allocated_gb": 2.5,
"memory_reserved_gb": 3.0
},
"model_loaded": true
}GET /api/
Access a simple web interface to generate images interactively.
| Variable | Description | Default |
|---|---|---|
MODEL_ID |
Hugging Face model ID | stabilityai/stable-diffusion-2-1-base |
AzureWebJobsStorage |
Storage connection string | Required |
FUNCTIONS_WORKER_RUNTIME |
Runtime identifier | python |
| Profile | GPU | vCPUs | Memory | Best For |
|---|---|---|---|---|
Consumption-GPU-NC8as-T4 |
NVIDIA T4 | 8 | 56 GB | Image generation, inference |
Consumption-GPU-NC16as-T4 |
NVIDIA T4 | 16 | 110 GB | Larger models |
Consumption-GPU-NC24as-T4 |
NVIDIA T4 | 24 | 220 GB | Multiple concurrent requests |
GPU workload profiles are available in:
- Sweden Central
- West US 3
- Australia East
- East US 2
Check the official documentation for the latest supported regions.
# Build the image
docker build -t gpu-image-gen:local -f Dockerfile .
# Run with GPU support
docker run --gpus all -p 7071:80 gpu-image-gen:local# Create virtual environment
python -m venv .venv
.venv\Scripts\activate # Windows
# or: source .venv/bin/activate # Linux/macOS
# Install dependencies (CPU-only PyTorch)
pip install -r requirements.txt
# Run locally
func start-
Reduce cold start time:
- Uncomment the model pre-download in Dockerfile
- Use artifact streaming (enable in Azure Portal)
- Keep
min-replicas >= 1for warm instances
-
Optimize inference:
- Enable xFormers for memory-efficient attention
- Use smaller image dimensions (512x512)
- Reduce inference steps (20-30 is usually sufficient)
-
Cost optimization:
- Set
min-replicas: 0when not in use - Use appropriate timeout values
- Monitor GPU utilization
- Set
- Azure Functions on Container Apps
- GPU Tutorial (Original Container)
- Serverless GPUs Overview
- Stable Diffusion on Hugging Face
To remove all resources:
az group delete --name gpu-functions-rg --yes --no-waitThis sample is provided under the MIT license.