From 9e2d0abdb0587ac8502a91cb66ad779c4ac48e94 Mon Sep 17 00:00:00 2001 From: Hou Chi Chan Date: Thu, 8 Oct 2026 10:11:18 -0700 Subject: [PATCH 01/14] Add offline multi-provider configuration validation Reuse setup profile validation for independent named provider selections, safe JSON results, and failing exit status. Add offline coverage and a focused usage guide without changing runtime routing. Co-authored-by: Copilot App <223556219+Copilot@users.noreply.github.com> Copilot-Session: 9314a109-37df-4021-8e75-bea995d57b23 --- .github/workflows/ci.yml | 3 + README.md | 3 + docs/MULTI-PROVIDER-VALIDATION.md | 133 ++++++++++++++++++++ setup/Test-EppProviders.ps1 | 31 +++++ setup/support/Epp.Setup.psm1 | 129 +++++++++++++++++--- setup/tests/Providers.Tests.ps1 | 196 ++++++++++++++++++++++++++++++ 6 files changed, 481 insertions(+), 14 deletions(-) create mode 100644 docs/MULTI-PROVIDER-VALIDATION.md create mode 100644 setup/Test-EppProviders.ps1 create mode 100644 setup/tests/Providers.Tests.ps1 diff --git a/.github/workflows/ci.yml b/.github/workflows/ci.yml index 7cd1695..e60d5aa 100644 --- a/.github/workflows/ci.yml +++ b/.github/workflows/ci.yml @@ -17,6 +17,9 @@ jobs: - name: Test certificate setup without Azure access shell: pwsh run: ./setup/tests/Certificates.Tests.ps1 + - name: Test provider validation without network or secret access + shell: pwsh + run: ./setup/tests/Providers.Tests.ps1 - name: Compile Bicep without deploying shell: pwsh run: | diff --git a/README.md b/README.md index 581f8b6..697e9ca 100644 --- a/README.md +++ b/README.md @@ -6,6 +6,9 @@ by SMS or voice. Start here to onboard **one deployment in one Azure region**. For implementation details, configuration, packaging, and security behavior, see the [technical reference](TECHNICAL.md). +To check multiple provider profiles locally without sending messages or changing a deployment, +see [offline multi-provider configuration validation](docs/MULTI-PROVIDER-VALIDATION.md). + ## Deployment options | Option | Onboarding | diff --git a/docs/MULTI-PROVIDER-VALIDATION.md b/docs/MULTI-PROVIDER-VALIDATION.md new file mode 100644 index 0000000..d330c13 --- /dev/null +++ b/docs/MULTI-PROVIDER-VALIDATION.md @@ -0,0 +1,133 @@ +# Offline multi-provider configuration validation + +Validate multiple named **setup provider profiles** in one local run before considering deployment. +This is not a Function invocation, SAS trigger, authenticated evaluation, or delivery test. +A passing result means **configuration validity only**: it does not prove credentials, provider +account entitlement, sender/channel approval, API availability, or handset delivery. + +## Prerequisites and coverage + +Use PowerShell 7+ and a local checkout of this repository containing `setup/support` and +`setup/providers`. Unlike the deployment launcher, the validator does not download missing files. +No Azure CLI, Graph modules, sign-in, Azure subscription access, Key Vault access, or provider +credentials are needed. Do not put secret values in profiles or command arguments. + +The shipped [setup catalog](../setup/providers/catalog.json) contains `telesign` and `soprano`. +Their [Telesign](../setup/providers/telesign.json) and [Soprano](../setup/providers/soprano.json) +profiles support `sms`/`voice` and `global`/`eu` selections. Provider scope is not an Azure region. +The Function runtimes also have Infobip and Sinch adapters, but **there are no setup profiles for +them in this catalog**; this command rejects those IDs rather than implying they are covered. +It does not validate arbitrary runtime app settings or discover providers from the environment. + +## Run from the repository root + +In PowerShell: + +```powershell +.\setup\Test-EppProviders.ps1 -Provider telesign,soprano -Channel sms -EndpointRegion global +$LASTEXITCODE +``` + +Validate the voice/EU selection, or retain a single-provider check: + +```powershell +.\setup\Test-EppProviders.ps1 -Provider telesign,soprano -Channel voice -EndpointRegion eu +.\setup\Test-EppProviders.ps1 -Provider telesign -Channel sms -EndpointRegion global +``` + +`Provider` is a PowerShell string array of catalog IDs. Case and surrounding whitespace are +normalized; aliases, unknown IDs, empty entries, and duplicates are rejected. Every occurrence of +a duplicate fails. The channel and provider scope are explicit and shared by the selections in +that run; run again for a different pair. They never change deployed routing. + +For a separate process (including CI), use `-Command` so PowerShell constructs the array; +do not pass a comma-separated string to `pwsh -File` and expect it to become an array: + +```powershell +pwsh -NoProfile -Command '& .\setup\Test-EppProviders.ps1 -Provider telesign,soprano -Channel sms -EndpointRegion global' +``` + +To check proposed profile edits, edit the local profile JSON or supply `-ProfileDirectory` pointing +to a local directory containing `catalog.json` and its profile JSON files: + +```powershell +.\setup\Test-EppProviders.ps1 -Provider telesign,soprano -Channel sms -EndpointRegion global ` + -ProfileDirectory .\candidate-providers +``` + +Keep the same catalog IDs and complete profile contracts when preparing candidates. Catalog +entries must have unique IDs and files and unambiguous display names. Passing a custom catalog +does not add runtime adapter support. No profile contents, resolved app settings, endpoint URLs, +tenant IDs, credential values, or raw parser exceptions are included in reports. + +## Output and failure behavior + +The launcher writes one JSON report and exits **0 only if all selected configurations pass**. +It exits **1** for any invalid selection, catalog, unreadable profile, or invalid profile. +An empty batch or an invalid/unreadable/ambiguous catalog fails the run with `Errors` populated +and no `Results`; no provider can be resolved safely in those cases. + +A successful two-provider report has this shape: + +```json +{ + "Valid": true, + "Errors": [], + "Results": [ + { "Index": 1, "Provider": "telesign", "Valid": true, "Code": "ConfigurationValid", "Issues": [] }, + { "Index": 2, "Provider": "soprano", "Valid": true, "Code": "ConfigurationValid", "Issues": [] } + ] +} +``` + +Results retain input order and a one-based `Index`. Unknown or empty provider selections have +`Provider: null` rather than echoing arbitrary input. Other result codes are `InvalidSelection`, +`DuplicateSelection`, `UnknownProvider`, `ProfileUnreadable`, and `ProfileInvalid`. `Issues` +contains safe descriptions; malformed input is not copied into an error message. + +For example, `-Provider telesign,sinch,soprano` returns three results: Telesign and Soprano can +pass independently, Sinch returns `UnknownProvider`, the overall `Valid` is `false`, and the +exit code is 1. Similarly, an invalid Telesign profile does not prevent Soprano from being checked. +Fix the reported profile fields or JSON syntax locally, then rerun the same command. + +For automation already running inside PowerShell, use the exported function without exiting +the caller; check `Valid` explicitly: + +```powershell +Import-Module .\setup\support\Epp.Setup.psm1 +$report = Test-EppProviderConfiguration -Provider telesign,soprano -Channel sms -EndpointRegion global +if (-not $report.Valid) { + $report | ConvertTo-Json -Depth 6 + throw 'Offline provider configuration validation failed.' +} +``` + +## Code and validation boundary + +- [CLI entry point](../setup/Test-EppProviders.ps1): JSON serialization and exit status. +- [Shared setup module](../setup/support/Epp.Setup.psm1): `Test-EppProviderConfiguration` resolves + independent selections using `Get-EppProviderCatalog`, then reuses + `ConvertTo-EppProviderSettings`, the same validator used by single-provider setup. +- [Focused tests](../setup/tests/Providers.Tests.ps1): independent results, invalid/duplicate + selections, safe errors, offline behavior, CLI status, and single-provider regression. + +The shared validator checks the enabled flag, provider identity, tenant GUID, authentication mode, +Key Vault **secret-name syntax** for API-key profiles, public HTTPS endpoint syntax, timeout/retry +bounds, and OAuth application ID/scope syntax. It requires **all four SMS/voice Global/EU routes**, +including routes not selected for this run, just as deployment setup does. Public hostname syntax +is not a DNS or reachability test. Secret names are configuration references, not fetched secrets. + +This command only reads local catalog/profile files. It performs no HTTP requests, provider +sends, credential reads/refresh/rotation, resource deployments, or Graph policy/role changes. +It does not contact SAS, call the Function, or rely on evaluation/fallback behavior. + +Existing deployment remains **one provider per endpoint**, selected by `EPP_PROVIDER_NAME`. +No runtime switching, fan-out, fallback, or multi-provider deployment is introduced. The +authenticated encrypted evaluation contract is unchanged: valid `mode: 2` requests return the +nonce before provider selection, credential resolution, or outbound provider calls. Offline +validation does not exercise or prove that contract on a deployed endpoint. Background credential +refresh in an existing deployment remains independent and is not triggered by this command. + +Deployment, authenticated evaluation, provider onboarding, and any separately authorized live +delivery checks remain distinct steps. Never interpret `ConfigurationValid` as permission to +send a message or as evidence that an exposed credential is safe to reuse. diff --git a/setup/Test-EppProviders.ps1 b/setup/Test-EppProviders.ps1 new file mode 100644 index 0000000..ab883f0 --- /dev/null +++ b/setup/Test-EppProviders.ps1 @@ -0,0 +1,31 @@ +#Requires -Version 7.0 +<# +.SYNOPSIS + Validate multiple local setup provider profiles without network or credential access. +.DESCRIPTION + Writes one JSON report. Exits 0 only when every selected provider configuration is valid; + otherwise exits 1. This does not test provider credentials, entitlement, or delivery. +.EXAMPLE + .\Test-EppProviders.ps1 -Provider telesign,soprano -Channel sms -EndpointRegion global +#> +[CmdletBinding()] +param( + [AllowEmptyCollection()] [AllowEmptyString()] [string[]] $Provider = @(), + [string] $Channel, + [string] $EndpointRegion, + [string] $ProfileDirectory = (Join-Path $PSScriptRoot 'providers') +) + +$ErrorActionPreference = 'Stop' +Set-StrictMode -Version Latest +$module = Import-Module (Join-Path $PSScriptRoot 'support/Epp.Setup.psm1') -PassThru -Force +try { + $report = Test-EppProviderConfiguration -Provider $Provider -Channel $Channel ` + -EndpointRegion $EndpointRegion -ProfileDirectory $ProfileDirectory + $report | ConvertTo-Json -Depth 6 + if (-not $report.Valid) { exit 1 } +} +finally { + Remove-Module -ModuleInfo $module +} +exit 0 diff --git a/setup/support/Epp.Setup.psm1 b/setup/support/Epp.Setup.psm1 index 981b847..80e0ca6 100644 --- a/setup/support/Epp.Setup.psm1 +++ b/setup/support/Epp.Setup.psm1 @@ -133,6 +133,37 @@ function Read-EppInput { } } +function Get-EppProviderCatalog { + param([string] $ProfileDirectory) + + $catalog = Read-EppJson (Join-Path $ProfileDirectory 'catalog.json') + if ($catalog['schemaVersion'] -ne 1 -or -not $catalog['providers']) { throw 'Unsupported or empty provider catalog.' } + $entries = @($catalog['providers']) + $ids = @{} + $aliases = @{} + $files = @{} + foreach ($entry in $entries) { + if ($entry -isnot [Collections.IDictionary] -or + $entry['id'] -isnot [string] -or $entry['id'] -cnotmatch '^[a-z][a-z0-9-]{1,31}$' -or + $entry['file'] -isnot [string] -or + $entry['file'] -cnotmatch '^[a-z][a-z0-9-]{1,31}\.json$' -or + $entry['displayName'] -isnot [string] -or [string]::IsNullOrWhiteSpace($entry['displayName']) -or + $entry['displayName'] -match '[\x00-\x1f]' -or $ids.ContainsKey($entry['id']) -or + $files.ContainsKey($entry['file'])) { + throw 'Provider catalog contains an invalid or duplicate entry.' + } + $ids[$entry['id']] = $true + $files[$entry['file']] = $true + foreach ($alias in @($entry['id'], $entry['displayName'])) { + if ($aliases.ContainsKey($alias) -and $aliases[$alias] -ne $entry['id']) { + throw 'Provider catalog contains an ambiguous ID or display name.' + } + $aliases[$alias] = $entry['id'] + } + } + return $entries +} + function Get-EppProvider { param( [string] $AssetDirectory, [string] $SourceBaseUri, [string] $Provider, [string] $Channel, @@ -145,18 +176,7 @@ function Get-EppProvider { if ($SourceBaseUri -cnotmatch $sourcePattern) { throw 'Provider files must come from the same commit-pinned selected repository as the deployment tools.' } - $catalog = Read-EppJson (Join-Path $AssetDirectory 'providers/catalog.json') - if ($catalog['schemaVersion'] -ne 1 -or -not $catalog['providers']) { throw 'Unsupported or empty provider catalog.' } - $entries = @($catalog['providers']) - $ids = @{} - foreach ($entry in $entries) { - if ($entry -isnot [Collections.IDictionary] -or $entry['id'] -cnotmatch '^[a-z][a-z0-9-]{1,31}$' -or - $entry['file'] -cnotmatch '^[a-z][a-z0-9-]{1,31}\.json$' -or - -not $entry['displayName'] -or $entry['displayName'] -match '[\x00-\x1f]' -or $ids.ContainsKey($entry['id'])) { - throw 'Provider catalog contains an invalid or duplicate entry.' - } - $ids[$entry['id']] = $true - } + $entries = @(Get-EppProviderCatalog -ProfileDirectory (Join-Path $AssetDirectory 'providers')) $selected = Select-EppOption -Entries $entries -Name Provider -Value $Provider -NonInteractive:$NonInteractive $path = Join-Path $AssetDirectory "providers/$($selected['file'])" @@ -236,7 +256,9 @@ function ConvertTo-EppProviderSettings { } } if ($issues.Count) { - throw "Provider '$DisplayName' is not deployment-ready:`n - $($issues -join "`n - ")`nAsk the provider owner to complete its GitHub JSON. No Azure resources were changed." + $validationError = [IO.InvalidDataException]::new("Provider '$DisplayName' is not deployment-ready:`n - $($issues -join "`n - ")`nAsk the provider owner to complete its GitHub JSON. No Azure resources were changed.") + $validationError.Data['EppValidationIssues'] = $issues.ToArray() + throw $validationError } $channelEntry = Select-EppOption -Entries @( @@ -273,6 +295,85 @@ function ConvertTo-EppProviderSettings { } } +function Test-EppProviderConfiguration { + [CmdletBinding()] + param( + [AllowEmptyCollection()] [AllowEmptyString()] [string[]] $Provider = @(), + [string] $Channel, + [string] $EndpointRegion, + [string] $ProfileDirectory = (Join-Path $PSScriptRoot '../providers') + ) + + $report = [pscustomobject]@{ Valid = $false; Errors = @(); Results = @() } + if ($null -eq $Provider -or -not $Provider.Count) { + $report.Errors = @('Select at least one provider catalog ID.') + return $report + } + try { $entries = @(Get-EppProviderCatalog -ProfileDirectory $ProfileDirectory) } + catch { + $report.Errors = @('Cannot read a valid, unambiguous local provider catalog. Check catalog.json and its entries.') + return $report + } + + $selections = @($Provider | ForEach-Object { ([string]$_).Trim().ToLowerInvariant() }) + $results = for ($index = 0; $index -lt $selections.Count; $index++) { + $id = $selections[$index] + $selected = $entries | Where-Object { $_['id'] -ceq $id } + $result = [pscustomobject]@{ + Index = $index + 1 + Provider = if ($selected) { $selected['id'] } else { $null } + Valid = $false + Code = 'InvalidSelection' + Issues = @() + } + if (-not $id) { + $result.Issues = @('Provider selection must not be empty.') + } + elseif (@($selections | Where-Object { $_ -ceq $id }).Count -gt 1) { + $result.Code = 'DuplicateSelection' + $result.Issues = @('Each provider catalog ID may be selected only once (case-insensitive).') + } + elseif (-not $selected) { + $result.Code = 'UnknownProvider' + $result.Issues = @('Provider selection must match an ID in the local setup catalog; aliases and runtime-only adapters are not supported.') + } + elseif ($Channel -notin @('sms', 'voice') -or $EndpointRegion -notin @('global', 'eu')) { + $result.Issues = @('Specify Channel sms or voice and EndpointRegion global or eu.') + } + else { + try { + $profile = Read-EppJson (Join-Path $ProfileDirectory $selected['file']) + } + catch { + $result.Code = 'ProfileUnreadable' + $result.Issues = @('Cannot read the local profile as a JSON object. Check the catalog file entry and JSON syntax.') + $result + continue + } + try { + $null = ConvertTo-EppProviderSettings -Profile $profile -Id $selected['id'] -DisplayName $selected['id'] ` + -Channel $Channel -EndpointRegion $EndpointRegion -NonInteractive + $result.Valid = $true + $result.Code = 'ConfigurationValid' + } + catch { + $result.Code = 'ProfileInvalid' + # Only validator-authored field descriptions are safe to return, not parser errors or input values. + $result.Issues = @( + if ($_.Exception.Data.Contains('EppValidationIssues')) { + $_.Exception.Data['EppValidationIssues'] + } + else { 'Profile structure is invalid. Check deployment, authentication, and all SMS/voice Global/EU routes.' } + ) + } + } + $result + } + $report.Results = @($results) + $report.Valid = @($report.Results | Where-Object { -not $_.Valid }).Count -eq 0 + return $report +} + function Get-EppResourceNames { param([string] $SubscriptionId, [string] $ApplicationId, [string] $ResourcePrefix) @@ -1684,4 +1785,4 @@ function Invoke-EppSetup { } } -Export-ModuleMember -Function Invoke-EppSetup +Export-ModuleMember -Function Invoke-EppSetup, Test-EppProviderConfiguration diff --git a/setup/tests/Providers.Tests.ps1 b/setup/tests/Providers.Tests.ps1 new file mode 100644 index 0000000..1b0962f --- /dev/null +++ b/setup/tests/Providers.Tests.ps1 @@ -0,0 +1,196 @@ +#Requires -Version 7.0 +$ErrorActionPreference = 'Stop' +Set-StrictMode -Version Latest +$module = Import-Module (Join-Path $PSScriptRoot '../support/Epp.Setup.psm1') -Force -PassThru +$directory = Join-Path ([IO.Path]::GetTempPath()) "epp-provider-tests-$([Guid]::NewGuid().ToString('N'))" +$profiles = Join-Path $directory 'providers' +New-Item -ItemType Directory -Path $profiles | Out-Null +Copy-Item (Join-Path $PSScriptRoot '../providers/*.json') -Destination $profiles +try { + & $module { + param($Directory, $Profiles) + + function script:Assert($Condition, [string] $Message) { + if (-not $Condition) { throw $Message } + } + $script:ForbiddenCalls = [Collections.Generic.List[string]]::new() + foreach ($command in @('Invoke-WebRequest', 'Invoke-RestMethod', 'Invoke-EppAz', 'Invoke-MgGraphRequest', + 'Get-AzKeyVaultSecret', 'Get-Secret', 'Read-Host')) { + Set-Item -Path "Function:script:$command" -Value { + $script:ForbiddenCalls.Add($MyInvocation.MyCommand.Name) + throw 'Network, credentials and prompts are forbidden in offline validation.' + } + } + $script:AllowedReads = @('catalog.json', 'telesign.json', 'soprano.json') | + ForEach-Object { Join-Path $Profiles $_ } + function script:Get-Content { + param($LiteralPath, [switch] $Raw) + Assert ($LiteralPath -in $script:AllowedReads) 'Validation must only read local catalog/profile JSON, never credentials.' + Microsoft.PowerShell.Management\Get-Content -LiteralPath $LiteralPath -Raw:$Raw + } + $parameters = @{ Provider = @('telesign', 'soprano'); Channel = 'sms'; EndpointRegion = 'global'; ProfileDirectory = $Profiles } + foreach ($channel in @('sms', 'voice')) { + foreach ($region in @('global', 'eu')) { + $report = Test-EppProviderConfiguration @parameters -Channel $channel -EndpointRegion $region + Assert ($report.Valid -and $report.Results.Count -eq 2 -and $report.Errors.Count -eq 0) 'Both shipped profiles must pass for every channel/region.' + Assert (($report.Results.Provider -join ',') -ceq 'telesign,soprano') 'Results must retain selection order and canonical provider IDs.' + Assert (@($report.Results | Where-Object { $_.Code -ne 'ConfigurationValid' -or $_.Issues.Count }).Count -eq 0) 'Successful results must contain no issues.' + } + } + $single = Test-EppProviderConfiguration @parameters -Provider ' TELESIGN ' + Assert ($single.Valid -and $single.Results.Count -eq 1 -and $single.Results[0].Provider -ceq 'telesign') 'Single selections must normalize case and surrounding whitespace.' + + foreach ($selection in @(@(), @(''), @(' '))) { + $report = Test-EppProviderConfiguration @parameters -Provider $selection + Assert (-not $report.Valid) 'Empty selection must fail without prompting or choosing a provider.' + } + $report = Test-EppProviderConfiguration @parameters -Provider $null + Assert (-not $report.Valid -and $report.Errors.Count -eq 1) 'Null selection must be an explicit failed report.' + $report = Test-EppProviderConfiguration @parameters -Provider @('telesign', '') + Assert ($report.Results[0].Valid -and $report.Results[1].Code -eq 'InvalidSelection' -and -not $report.Valid) 'An empty entry must not prevent another provider from being validated.' + $report = Test-EppProviderConfiguration @parameters -Provider @('telesign', ' TELESIGN ', 'soprano') + Assert (-not $report.Valid -and $report.Results.Count -eq 3) 'Duplicate selections must fail the run.' + Assert ($report.Results[0].Code -eq 'DuplicateSelection' -and $report.Results[1].Code -eq 'DuplicateSelection' -and $report.Results[2].Valid) 'Every occurrence of a duplicate must fail, without skipping independent providers.' + $report = Test-EppProviderConfiguration @parameters -Provider @('infobip', 'sinch', 'soprano', 'not-a-provider') + Assert (-not $report.Valid -and $report.Results[2].Valid) 'Runtime-only and unknown providers must fail without stopping known providers.' + Assert (@($report.Results | Where-Object Code -eq 'UnknownProvider').Count -eq 3) 'Unknown selections must be identified explicitly.' + foreach ($invalidRoute in @(@{ Channel = '' }, @{ Channel = 'fax' }, @{ EndpointRegion = '' }, @{ EndpointRegion = 'westus2' })) { + $arguments = $parameters.Clone() + foreach ($key in $invalidRoute.Keys) { $arguments[$key] = $invalidRoute[$key] } + $report = Test-EppProviderConfiguration @arguments + Assert (-not $report.Valid -and @($report.Results | Where-Object Code -eq 'InvalidSelection').Count -eq 2) 'Invalid/missing routes must not silently select a default.' + } + + $telesignPath = Join-Path $Profiles 'telesign.json' + $original = Get-Content -LiteralPath $telesignPath -Raw + $catalogPath = Join-Path $Profiles 'catalog.json' + $catalogOriginal = Get-Content -LiteralPath $catalogPath -Raw + $sentinel = 'DO-NOT-REPORT-PRIVATE-INPUT' + $profile = $original | ConvertFrom-Json -AsHashtable + $profile.deployment.routes.voice.eu.endpoint = "https://user:$sentinel@provider.contoso.com/?token=$sentinel" + $profile.deployment.routes.sms.global.timeoutMilliseconds = 2501 + $profile.deployment.routes.sms.global.retryIntervalSeconds = -1 + $profile.deployment.authentication.keyVaultSecretName = $sentinel + $profile | ConvertTo-Json -Depth 10 | Set-Content -LiteralPath $telesignPath + $report = Test-EppProviderConfiguration @parameters + Assert (-not $report.Valid -and -not $report.Results[0].Valid -and $report.Results[1].Valid) 'A bad profile must not stop validation of a good profile.' + Assert ($report.Results[0].Code -eq 'ProfileInvalid') 'Invalid profiles must be reported explicitly.' + $json = $report | ConvertTo-Json -Depth 6 + Assert (-not $json.Contains($sentinel)) 'Reports must never echo invalid profile values.' + Assert ($json.Contains('deployment.routes.voice.eu.endpoint') -and $json.Contains('timeoutMilliseconds') -and $json.Contains('retryIntervalSeconds') -and $json.Contains('keyVaultSecretName')) 'Safe issue descriptions must identify invalid fields, including unselected routes.' + + foreach ($badJson in @("{ malformed-$sentinel", 'null', '[]')) { + Set-Content -LiteralPath $telesignPath -Value $badJson + $report = Test-EppProviderConfiguration @parameters + Assert (-not $report.Valid -and $report.Results[0].Code -eq 'ProfileUnreadable' -and $report.Results[1].Valid) 'Malformed/nonobject JSON must fail independently.' + Assert (-not ($report | ConvertTo-Json -Depth 6).Contains($sentinel)) 'Parser exception details must not leak input values.' + } + foreach ($badProfile in @('{}', '{"deployment":{}}', '{"deployment":{"routes":[]}}')) { + Set-Content -LiteralPath $telesignPath -Value $badProfile + $report = Test-EppProviderConfiguration @parameters + Assert ($report.Results[0].Code -eq 'ProfileInvalid' -and $report.Results[1].Valid) 'Missing profile structure must be a safe independent failure.' + Assert ($report.Results[0].Issues -is [array]) 'Even a single structural issue must serialize as a JSON array.' + } + Remove-Item -LiteralPath $telesignPath + $report = Test-EppProviderConfiguration @parameters + Assert ($report.Results[0].Code -eq 'ProfileUnreadable' -and $report.Results[1].Valid) 'A missing profile must not trigger a download.' + Set-Content -LiteralPath $telesignPath -Value $original + foreach ($mutation in @('disabled', 'identity', 'tenant', 'authentication', 'timeout', 'retry')) { + $profile = $original | ConvertFrom-Json -AsHashtable + switch ($mutation) { + 'disabled' { $profile.deployment.enabled = $false } + 'identity' { $profile.deployment.providerName = 'soprano' } + 'tenant' { $profile.deployment.tenantId = $sentinel } + 'authentication' { $profile.deployment.authentication.mode = $sentinel } + 'timeout' { $profile.deployment.routes.sms.global.timeoutMilliseconds = 0 } + 'retry' { $profile.deployment.routes.sms.global.retryIntervalSeconds = 2147484 } + } + $profile | ConvertTo-Json -Depth 10 | Set-Content -LiteralPath $telesignPath + $report = Test-EppProviderConfiguration @parameters + Assert ($report.Results[0].Code -eq 'ProfileInvalid' -and $report.Results[1].Valid) "Shared validator must reject $mutation independently." + Assert ($report.Results[0].Issues -is [array] -and $report.Results[0].Issues.Count -eq 1) 'Single field issues must retain array shape.' + Assert (-not ($report | ConvertTo-Json -Depth 6).Contains($sentinel)) 'Invalid tenant/authentication values must not leak.' + } + Set-Content -LiteralPath $telesignPath -Value $original + $sopranoPath = Join-Path $Profiles 'soprano.json' + $sopranoOriginal = Get-Content -LiteralPath $sopranoPath -Raw + $profile = $sopranoOriginal | ConvertFrom-Json -AsHashtable + $profile.deployment.routes.voice.eu.appId = $sentinel + $profile.deployment.routes.voice.eu.scope = $sentinel + $profile | ConvertTo-Json -Depth 10 | Set-Content -LiteralPath $sopranoPath + $report = Test-EppProviderConfiguration @parameters + Assert (-not $report.Valid -and $report.Results[0].Valid -and $report.Results[1].Code -eq 'ProfileInvalid') 'Invalid OAuth settings must fail independently, including unselected routes.' + Assert ($report.Results[1].Issues.Count -eq 2 -and -not ($report | ConvertTo-Json -Depth 6).Contains($sentinel)) 'OAuth issues must use safe field descriptions.' + Set-Content -LiteralPath $sopranoPath -Value $sopranoOriginal + + foreach ($mutation in @('duplicate-id', 'duplicate-file', 'ambiguous-alias', 'path-traversal', 'schema')) { + $catalog = $catalogOriginal | ConvertFrom-Json -AsHashtable + switch ($mutation) { + 'duplicate-id' { $catalog.providers[1].id = 'telesign' } + 'duplicate-file' { $catalog.providers[1].file = 'telesign.json' } + 'ambiguous-alias' { $catalog.providers[1].displayName = 'Telesign' } + 'path-traversal' { $catalog.providers[0].file = '../private.json' } + 'schema' { $catalog.schemaVersion = 2 } + } + $catalog | ConvertTo-Json -Depth 6 | Set-Content -LiteralPath $catalogPath + $report = Test-EppProviderConfiguration @parameters + Assert (-not $report.Valid -and $report.Results.Count -eq 0 -and $report.Errors.Count -eq 1) 'Invalid or ambiguous catalogs must fail the entire run before profile reads.' + } + Set-Content -LiteralPath $catalogPath -Value "{ malformed-$sentinel" + $report = Test-EppProviderConfiguration @parameters + Assert (-not $report.Valid -and -not ($report | ConvertTo-Json -Depth 6).Contains($sentinel)) 'Catalog parser failures must not leak input.' + Set-Content -LiteralPath $catalogPath -Value $catalogOriginal + $report = Test-EppProviderConfiguration @parameters -Provider @($sentinel, 'soprano') + Assert (-not $report.Valid -and $report.Results[0].Index -eq 1 -and $null -eq $report.Results[0].Provider -and $report.Results[1].Valid) 'Unknown selections must be identified by position rather than echoed.' + Assert (-not ($report | ConvertTo-Json -Depth 6).Contains($sentinel)) 'Unknown selections must not leak arbitrary input.' + Assert ($script:ForbiddenCalls.Count -eq 0) 'Offline validation must not attempt network, secret access, or prompts even on failures.' + + $script:Downloads = @() + function script:Invoke-WebRequest { + param($Uri, $OutFile, $TimeoutSec, $MaximumRedirection) + $script:Downloads += $Uri + } + foreach ($id in @('telesign', 'soprano')) { + $selected = Get-EppProvider -AssetDirectory $Directory ` + -SourceBaseUri "https://raw.githubusercontent.com/Azure-Samples/ExternalPhoneProvider-AzureFunction-Sample/$('a' * 40)/setup" ` + -Provider $id -Channel voice -EndpointRegion eu -NonInteractive + $expected = ConvertTo-EppProviderSettings -Profile (Read-EppJson (Join-Path $Profiles "$id.json")) ` + -Id $id -DisplayName $id -Channel voice -EndpointRegion eu -NonInteractive + foreach ($key in $expected.Settings.Keys) { + Assert ($selected.Settings[$key] -ceq $expected.Settings[$key]) 'Single-provider setup must preserve its existing app settings.' + } + Assert ($selected.Settings.EPP_PROVIDER_NAME -ceq $id -and $selected.Channel -eq 'voice' -and $selected.EndpointRegion -eq 'eu') 'Setup must still select exactly one provider and route.' + } + Assert ($script:Downloads.Count -eq 2) 'Existing setup must retain one pinned profile download per selection.' + } $directory $profiles + + $launcher = (Join-Path $PSScriptRoot '../Test-EppProviders.ps1').Replace("'", "''") + $escapedProfiles = $profiles.Replace("'", "''") + foreach ($case in @( + @{ Selection = 'telesign,soprano'; Exit = 0; Count = 2 }, + @{ Selection = 'telesign'; Exit = 0; Count = 1 }, + @{ Selection = 'telesign,unknown'; Exit = 1; Count = 2 }, + @{ Selection = 'telesign,TELESIGN'; Exit = 1; Count = 2 }, + @{ Selection = '@()'; Exit = 1; Count = 0 }, + @{ Selection = 'telesign,soprano'; Exit = 1; Count = 2; InvalidProfile = $true } + )) { + if ($case.ContainsKey('InvalidProfile')) { + $path = Join-Path $profiles 'telesign.json' + $profile = Get-Content -LiteralPath $path -Raw | ConvertFrom-Json -AsHashtable + $profile.deployment.enabled = $false + $profile | ConvertTo-Json -Depth 10 | Set-Content -LiteralPath $path + } + $output = & (Join-Path $PSHOME 'pwsh') -NoProfile -Command "& '$launcher' -Provider $($case.Selection) -Channel sms -EndpointRegion global -ProfileDirectory '$escapedProfiles'" 2>&1 + if ($LASTEXITCODE -ne $case.Exit) { throw "Unexpected CLI exit code for $($case.Selection): $LASTEXITCODE" } + $report = ($output -join "`n") | ConvertFrom-Json + if ($report.Valid -ne ($case.Exit -eq 0) -or $report.Results.Count -ne $case.Count) { throw 'CLI must emit exactly one JSON report matching its exit code.' } + foreach ($result in $report.Results) { + if ($result.Issues -isnot [array]) { throw 'CLI issues must always be JSON arrays, including single-issue failures.' } + } + } + Write-Host 'Provider validation tests passed (offline batch, safe failures, CLI exit codes, single-provider regression).' +} +finally { + Remove-Module -ModuleInfo $module + Remove-Item -LiteralPath $directory -Recurse -Force +} From 3f34fe77aa39659945f5f3b559150ea634da4814 Mon Sep 17 00:00:00 2001 From: Hou Chi Chan Date: Thu, 8 Oct 2026 10:31:12 -0700 Subject: [PATCH 02/14] Replace validation feature with customer implementation guide Remove the shipped setup validator, tests and CI additions. Document offline runtime feasibility, isolated provider contexts, existing code seams and customer-owned rollout requirements without adding a router. Co-authored-by: Copilot App <223556219+Copilot@users.noreply.github.com> Copilot-Session: 9314a109-37df-4021-8e75-bea995d57b23 --- .github/workflows/ci.yml | 3 - README.md | 4 +- docs/MULTI-PROVIDER-IMPLEMENTATION.md | 283 ++++++++++++++++++++++++++ docs/MULTI-PROVIDER-VALIDATION.md | 133 ------------ setup/Test-EppProviders.ps1 | 31 --- setup/support/Epp.Setup.psm1 | 129 ++---------- setup/tests/Providers.Tests.ps1 | 196 ------------------ 7 files changed, 299 insertions(+), 480 deletions(-) create mode 100644 docs/MULTI-PROVIDER-IMPLEMENTATION.md delete mode 100644 docs/MULTI-PROVIDER-VALIDATION.md delete mode 100644 setup/Test-EppProviders.ps1 delete mode 100644 setup/tests/Providers.Tests.ps1 diff --git a/.github/workflows/ci.yml b/.github/workflows/ci.yml index e60d5aa..7cd1695 100644 --- a/.github/workflows/ci.yml +++ b/.github/workflows/ci.yml @@ -17,9 +17,6 @@ jobs: - name: Test certificate setup without Azure access shell: pwsh run: ./setup/tests/Certificates.Tests.ps1 - - name: Test provider validation without network or secret access - shell: pwsh - run: ./setup/tests/Providers.Tests.ps1 - name: Compile Bicep without deploying shell: pwsh run: | diff --git a/README.md b/README.md index 697e9ca..6f74eea 100644 --- a/README.md +++ b/README.md @@ -6,8 +6,8 @@ by SMS or voice. Start here to onboard **one deployment in one Azure region**. For implementation details, configuration, packaging, and security behavior, see the [technical reference](TECHNICAL.md). -To check multiple provider profiles locally without sending messages or changing a deployment, -see [offline multi-provider configuration validation](docs/MULTI-PROVIDER-VALIDATION.md). +For a customer-built multi-provider customization, see the +[implementation guide and offline feasibility findings](docs/MULTI-PROVIDER-IMPLEMENTATION.md). ## Deployment options diff --git a/docs/MULTI-PROVIDER-IMPLEMENTATION.md b/docs/MULTI-PROVIDER-IMPLEMENTATION.md new file mode 100644 index 0000000..b5e31c3 --- /dev/null +++ b/docs/MULTI-PROVIDER-IMPLEMENTATION.md @@ -0,0 +1,283 @@ +# Implementing multiple provider configurations + +**A customer customization is feasible, but it is not a built-in feature.** An offline JavaScript +prototype exercised the existing adapters with independently configured Telesign API-key and +Soprano OAuth contexts, including concurrent requests. The critical requirement is to isolate +configuration and credentials per context, not merely change the selected provider name. + +This guide describes changes **you would implement in your own application**. It does not ship a +router, setup script, configuration validator, or new deployment behavior. The unmodified sample +still uses **one provider per deployment**, selected by `EPP_PROVIDER_NAME`. Separate deployments +remain the simpler option when accounts require strong isolation. + +## What was demonstrated offline + +A session-only experiment against the JavaScript implementation at commit +`0d5db07f00a8b83295d83344a877f3bfd1ef009a`, using Node.js 22.17.1, passed six focused tests: + +- Forty concurrent synthetic SMS dispatches used four named contexts: two Telesign accounts and + two Soprano accounts. Each had its own frozen configuration and `CredentialTokenService`. + Assertions checked the exact endpoint, account/token, request content, correlation ID and + response mapping for every dispatch. Concurrent credential acquisition coalesced within each + context, not across accounts. +- A negative control deliberately reused one credential service with another account and then + another authentication mode. It returned the first cached API-key bundle. **Changing the + arguments passed to the existing singleton does not switch its configuration.** +- A locally encrypted evaluation request was parsed and decrypted, then returned its nonce + before route lookup, credential acquisition or transport, even with an unknown route ID. + An incomplete decrypted context still failed before routing. +- Unknown/duplicate route IDs, an unknown adapter and a channel mismatch were rejected before + credential or transport calls. +- A synthetic provider HTTP failure was terminal for that request while an independent + concurrent route succeeded; no retry or alternate-provider dispatch occurred. +- After simulated secret-cache expiry, a failing account could not borrow another context's + credentials or cause fallback. An independent OAuth route still succeeded. + +The prototype used the **real** configuration reader, payload parser, local JWE decryption, +credential service/cache classes, Telesign/Soprano request builders, shared transport logic and +response mappers. Only the Azure SDK boundary and `fetch` were replaced with deterministic fakes; +socket, HTTP and DNS access were guarded to fail. Periodic refresh scheduling was simulated and +every service was closed. No process environment was changed per request. Existing JavaScript +adapter, credential-cache and handler tests also passed. + +**This was not an end-to-end provider or SAS test.** The prototype used a customer-style dispatch +seam, test-owned route IDs, fake credentials, fake responses and a locally generated encryption key. +It did not run a deployed Function, validate Easy Auth, establish a trusted routing identity, +perform a real OAuth exchange, read Key Vault, or send SMS/voice messages. Its new multi-context +dispatch tests covered SMS, not a multi-provider voice rollout. Real refresh scheduling, distributed +workers, production load and configuration reload were not demonstrated. The prototype is not +included as a supported repository command; the evidence does not certify a production router. + +## Start at these existing code seams + +### JavaScript: experimentally exercised + +- [SendOtp.js](../javascript/src/functions/SendOtp.js): the registered handler validates/decrypts, + returns evaluation before selection, checks channel/authentication/endpoint, resolves credentials, + builds one request, calls transport and maps the endpoint response. Its + `startProviderCredentialRefresh`/`stopProviderCredentialRefresh` hooks currently manage one + selected credential service. +- [config.js](../javascript/src/functions/config.js): `readConfig(env)` / `AppConfig` already accept + an explicit settings object; use that seam rather than mutating `process.env`. +- [providers/index.js](../javascript/src/functions/providers/index.js): `selectProvider(name)` + resolves a fixed adapter with no default. [Telesign](../javascript/src/functions/providers/telesign.js) + and [Soprano](../javascript/src/functions/providers/soprano.js) expose `credentialSpec`, + `createRequest` and `interpretResponse`. +- [credentials.js](../javascript/src/functions/credentials.js): `CredentialTokenService`, + `ApiKeyCache` and `AccessTokenCache` own acquisition, cache expiry, in-flight refresh and shutdown. + The exported singleton owns **one** selected cache; it is not keyed by provider or configuration. +- [providerTransport.js](../javascript/src/functions/providerTransport.js): `sendProviderRequest` + enforces request URL checks, timeout and manual redirects. + [providerResult.js](../javascript/src/functions/providerResult.js) maps adapter outcomes to + endpoint statuses; preserve those mappings rather than treating all HTTP responses as success. +- [entraPayload.js](../javascript/src/functions/entraPayload.js), + [jwe.js](../javascript/src/functions/jwe.js) and + [logging.js](../javascript/src/functions/logging.js): retain payload/decryption checks and + selected-metadata logging. Parsing a `tenantId` field does not establish its routing authority. + +### .NET: inspected, not multi-context-tested here + +Follow [Functions/SendOtp.cs](../dotnet/Functions/SendOtp.cs) for evaluation, `SelectProvider` and +dispatch; [AppConfig.Read(IEnv)](../dotnet/Src/AppConfig.cs) and +[IEnv](../dotnet/Src/Models.cs) provide explicit configuration seams. +[PhoneProviderBase](../dotnet/Src/PhoneProviderBase.cs) defines credential acquisition, +`SendOtpAsync`, shared HTTP handling and response-status mapping. + +[CredentialTokenService](../dotnet/Src/CredentialTokenService.cs) caches by `provider.Name`, which +does not distinguish two accounts of the same adapter. +[SopranoProvider](../dotnet/Src/Providers/SopranoProvider.cs) also retains the first initialized +OAuth identity/credential/scope. [SecretResolver](../dotnet/Src/SecretResolver.cs) retains a vault +client built from its `IEnv`. Merely changing the outer cache key is therefore insufficient. +Build a context-owned service/provider/resolver/configuration object graph, or redesign every +relevant cache key and initialization boundary. Review cold-acquisition concurrency explicitly; +do not assume its cache has the same single-flight behavior as JavaScript. +[Program.cs](../dotnet/Program.cs) currently registers singleton providers/services and a +non-redirecting HTTP client; customer DI/lifecycle changes must preserve transport safety. + +### Python: inspected, not multi-context-tested here + +In [function_app.py](../python/function_app.py), `_select_provider`, `_send_to_provider` and the +evaluation branch are the dispatch seams. `_send_to_provider` currently rereads process settings. +Pass the selected context explicitly instead. [read_config(env)](../python/src/config.py) already +accepts a mapping. [CredentialTokenService](../python/src/credentials.py) owns one selected cache; +construct one per context, with a context-specific [SecretResolver(env)](../python/src/secrets.py). +Do not share the module-level `_credentials` across configurations. Update the warmup thread and +shutdown lifecycle deliberately. +[PhoneProviderBase](../python/src/provider.py) provides `build_request`, `map_response`, +`send_otp`, transport handling and endpoint-status mapping. + +The runtime adapters include Infobip, Sinch, Soprano and Telesign, but this experiment exercised +only Telesign and Soprano. The [setup catalog](../setup/providers/catalog.json) provisions only +Telesign/Soprano profiles. Adapter availability, setup coverage and a provider's approved account +capabilities are different things; none establishes a new multi-provider deployment contract. + +## Design your customization + +### Define named contexts and a trusted, deterministic selection rule + +A context ID identifies a complete configuration, not just an adapter name. For example, +`telesign-primary` and `telesign-secondary` must remain distinct even though both use Telesign. +Create a finite, reviewed context list at startup. Reject duplicate IDs, unknown adapters, +ambiguous rules, missing credentials configuration, disallowed endpoints and mismatched +authentication/channel settings before enabling live use. + +The following is **illustrative customer-owned configuration, not a supported new file format**. +Values in angle brackets must be supplied and reviewed; never place secret values in it: + +```json +[ + { + "id": "telesign-primary", + "settings": { + "EPP_PROVIDER_NAME": "telesign", + "EPP_PROVIDER_CHANNEL": "sms", + "EPP_PROVIDER_AUTH_MODE": "apiKey", + "EPP_PROVIDER_ENDPOINT": "https:///", + "EPP_PROVIDER_TIMEOUT_MS": "1500", + "KEY_VAULT_URL": "https://.vault.azure.net", + "AZURE_CLIENT_ID": "" + } + }, + { + "id": "soprano-primary", + "settings": { + "EPP_PROVIDER_NAME": "soprano", + "EPP_PROVIDER_CHANNEL": "voice", + "EPP_PROVIDER_AUTH_MODE": "oauth", + "EPP_PROVIDER_ENDPOINT": "https:///", + "EPP_PROVIDER_TIMEOUT_MS": "1500", + "EPP_PROVIDER_TENANT_ID": "", + "EPP_PROVIDER_SCOPE": "api:///.default", + "EPP_OUTBOUND_CLIENT_ID": "", + "EPP_OUTBOUND_MI_CLIENT_ID": "" + } + } +] +``` + +For one endpoint, an explicit server-owned `sms -> telesign-primary`, +`voice -> soprano-primary` rule is one possible deterministic policy. Validate the incoming channel +using the existing contract and authorize it against that policy. This is a design example; the +experiment used test-supplied route IDs for SMS across all four contexts. + +If different customers/accounts need the same channel, first define and verify the **trusted +endpoint/customer binding** from which routing can be authorized. Do not assume the SAS caller +identity distinguishes every customer. Do not use an arbitrary request header, a newly invented +`providerId`, correlation ID, or unverified body `tenantId` as authority. The existing wire contract +does not supply a new trusted provider-selector field. Never allow a caller to supply an endpoint, +vault, scope or secret name. Unknown or ambiguous routing must fail closed without a default. + +### Keep state and credentials isolated + +Build each context from a copied, immutable settings object. In JavaScript, `readConfig` retains +the supplied `env` reference, so freeze that object as well as the returned config. Separate +inbound authentication/decryption settings from outbound provider settings; do not switch the +inbound key because a provider route changed. + +Create **one long-lived credential service per context**, not per request. Bind it permanently to +that context's adapter credential specification, vault/secret names, vault-reading identity, +OAuth tenant/resource scope, outbound application and outbound managed identity. A provider name +alone is not a safe cache key. Share nothing that can retain another account's token or secret. +For Telesign, the existing adapter uses `telesign-api-key` and `telesign-customer-id`; distinct vaults +can isolate accounts with those same names. Using different names in one vault requires an explicit +per-context credential-spec customization, not mutation of the shared adapter object. + +Keep expiry, acquisition bounds, failure reporting and refresh coalescing. A failed/expired +context must fail its own request, never borrow credentials. Bound the context count, close every +service at shutdown, and define how background refresh is started. A request's evaluation branch +must not resolve credentials; independently running warmup/refresh is a separate lifecycle concern. +For configuration changes, prefer a controlled worker restart or a versioned replacement of the +whole context with safe draining/disposal, rather than modifying a live cache's inputs. + +These are logical isolation boundaries, not tenant security boundaries: one compromised worker +may access all identities assigned to it. Review least-privilege identity/vault access and provider +onboarding separately; retain separate deployments where stronger isolation is required. + +### Insert selection after evaluation, then dispatch once + +The sketch below uses existing JavaScript interfaces but is **pseudocode, not a complete handler**. +The uppercase helpers are customer responsibilities; authorization, error handling, lifecycle and +response/logging behavior must be integrated with the existing handler rather than bypassed. + +```javascript +// Startup: customer validates unique IDs, rules, endpoints, identities and required settings. +const contexts = new Map(); +for (const entry of VALIDATED_CUSTOMER_CONTEXTS) { + const settings = Object.freeze({ ...entry.settings }); + const config = Object.freeze(readConfig(settings)); + const provider = selectProvider(config.providerName); + REQUIRE_VALID_CONTEXT(entry.id, provider, config, contexts); + contexts.set(entry.id, Object.freeze({ + config, provider, credentials: new CredentialTokenService() + })); +} + +// Request: retain Easy Auth, envelope validation and JWE checks from SendOtp. +const { payload, deliveryContext } = VALIDATE_AND_DECRYPT_WITH_EXISTING_HANDLER(request); +if (payload.isEvaluation) { + return EXISTING_EVALUATION_RESPONSE(deliveryContext.nonce); +} +const id = AUTHORIZE_AND_SELECT_ONE_CONTEXT(serverOwnedPolicy, payload.channelName); +const context = contexts.get(id); +REQUIRE_KNOWN_CONTEXT_AND_MATCHING_CHANNEL(context, payload.channelName); +const { config, provider, credentials } = context; +const credential = await credentials.getCredentials(provider.credentialSpec, config); +REQUIRE_COMPLETE_CREDENTIAL_FOR_ADAPTER(credential, provider); +const delivery = BUILD_EXISTING_OTP_DELIVERY(deliveryContext, payload); +const outbound = provider.createRequest({ + channel: payload.channelName, endpoint: config.providerEndpoint, + delivery, credential, env: config.env +}); +const response = await sendProviderRequest( + outbound, parseProviderTimeout(config.providerTimeoutMs), safeLogContext); +const result = provider.interpretResponse(response); +return EXISTING_ENDPOINT_RESPONSE_MAPPING(result, deliveryContext.nonce); +// Shutdown: close each context.credentials; never fall through to another context on failure. +``` + +Keep approved endpoint allowlists in addition to generic HTTPS checks, preserve redirect/timeout +controls, and never forward inbound authorization to a provider. Retain the existing success, +block, failure and nonce rules. Do not turn an unknown provider response into success. +Log an approved non-secret context identifier and fixed failure classification, not request bodies, +OTP/phone data, credentials, tokens, raw provider descriptions or exception contents. +This design chooses **one** context per request: no fan-out, automatic fallback or resend. + +## Validate your implementation and roll out separately + +1. Start with offline fixtures. Adapt the real interfaces in + [provider-flow.test.js](../javascript/test/provider-flow.test.js), + [credential-cache.test.js](../javascript/test/credential-cache.test.js) and + [sendotp.test.js](../javascript/test/sendotp.test.js). + Replace SDK acquisition and transport with fakes, guard network access, and use only synthetic + delivery data. Test both authentication modes and multiple accounts of the same adapter under + concurrency; assert endpoint, credentials and correlation isolation, not just success counts. +2. Add unknown/duplicate/ambiguous routes, invalid settings, channel mismatch, cross-account + attempts, failed/expired credential acquisition, provider rejection, timeout, malformed response, + shutdown and configuration replacement tests. Prove failures cause no second provider attempt. + Verify valid evaluation returns before routing/credentials and malformed evaluation still fails. + Extend voice and every additional adapter before claiming support. +3. Run your runtime's existing contracts as well. .NET examples are + [SendOtpTests](../dotnet/tests/SendOtpTests.cs), + [CredentialTokenServiceTests](../dotnet/tests/CredentialTokenServiceTests.cs) and + [ContractTests](../dotnet/tests/ContractTests.cs). Python examples are + [test_engine.py](../python/tests/test_engine.py), + [test_credential_cache.py](../python/tests/test_credential_cache.py), + [test_function_app.py](../python/tests/test_function_app.py) and + [test_contract.py](../python/tests/test_contract.py). The shared + [contract fixtures](../tests/fixtures/contract.json) help preserve response behavior. +4. Review identity permissions, provider entitlement/consent, route and sender approval, secret + lifecycle, authorization boundaries, rate limits, observability and operational ownership. + Resolve exposed-credential rotation before any live validation. Passing a local configuration + or mocked test establishes none of these. +5. Deploy only after separate approval to a dedicated nonproduction environment. Verify deployed + authentication, encryption, policy readback and authorized evaluation independently. A real SAS + trigger can have surrounding fallback behavior: do not assume evaluation is harmless merely + because this Function returns before its provider call. +6. Perform live provider/delivery validation only with explicit authorization, approved recipient, + coordinated policy/attempt limits and a one-attempt safety gate. Provider acceptance is not + proof of handset delivery. Use an explicit rollout/rollback decision; do not add automatic + fallback to conceal a failing integration. + +No Azure resources, Graph roles/policy, Key Vault credentials, SAS triggers or provider sends were +used or changed to produce this guide. The feasibility result concerns reusable code seams and +offline context isolation, not production readiness or a delivered multi-provider feature. diff --git a/docs/MULTI-PROVIDER-VALIDATION.md b/docs/MULTI-PROVIDER-VALIDATION.md deleted file mode 100644 index d330c13..0000000 --- a/docs/MULTI-PROVIDER-VALIDATION.md +++ /dev/null @@ -1,133 +0,0 @@ -# Offline multi-provider configuration validation - -Validate multiple named **setup provider profiles** in one local run before considering deployment. -This is not a Function invocation, SAS trigger, authenticated evaluation, or delivery test. -A passing result means **configuration validity only**: it does not prove credentials, provider -account entitlement, sender/channel approval, API availability, or handset delivery. - -## Prerequisites and coverage - -Use PowerShell 7+ and a local checkout of this repository containing `setup/support` and -`setup/providers`. Unlike the deployment launcher, the validator does not download missing files. -No Azure CLI, Graph modules, sign-in, Azure subscription access, Key Vault access, or provider -credentials are needed. Do not put secret values in profiles or command arguments. - -The shipped [setup catalog](../setup/providers/catalog.json) contains `telesign` and `soprano`. -Their [Telesign](../setup/providers/telesign.json) and [Soprano](../setup/providers/soprano.json) -profiles support `sms`/`voice` and `global`/`eu` selections. Provider scope is not an Azure region. -The Function runtimes also have Infobip and Sinch adapters, but **there are no setup profiles for -them in this catalog**; this command rejects those IDs rather than implying they are covered. -It does not validate arbitrary runtime app settings or discover providers from the environment. - -## Run from the repository root - -In PowerShell: - -```powershell -.\setup\Test-EppProviders.ps1 -Provider telesign,soprano -Channel sms -EndpointRegion global -$LASTEXITCODE -``` - -Validate the voice/EU selection, or retain a single-provider check: - -```powershell -.\setup\Test-EppProviders.ps1 -Provider telesign,soprano -Channel voice -EndpointRegion eu -.\setup\Test-EppProviders.ps1 -Provider telesign -Channel sms -EndpointRegion global -``` - -`Provider` is a PowerShell string array of catalog IDs. Case and surrounding whitespace are -normalized; aliases, unknown IDs, empty entries, and duplicates are rejected. Every occurrence of -a duplicate fails. The channel and provider scope are explicit and shared by the selections in -that run; run again for a different pair. They never change deployed routing. - -For a separate process (including CI), use `-Command` so PowerShell constructs the array; -do not pass a comma-separated string to `pwsh -File` and expect it to become an array: - -```powershell -pwsh -NoProfile -Command '& .\setup\Test-EppProviders.ps1 -Provider telesign,soprano -Channel sms -EndpointRegion global' -``` - -To check proposed profile edits, edit the local profile JSON or supply `-ProfileDirectory` pointing -to a local directory containing `catalog.json` and its profile JSON files: - -```powershell -.\setup\Test-EppProviders.ps1 -Provider telesign,soprano -Channel sms -EndpointRegion global ` - -ProfileDirectory .\candidate-providers -``` - -Keep the same catalog IDs and complete profile contracts when preparing candidates. Catalog -entries must have unique IDs and files and unambiguous display names. Passing a custom catalog -does not add runtime adapter support. No profile contents, resolved app settings, endpoint URLs, -tenant IDs, credential values, or raw parser exceptions are included in reports. - -## Output and failure behavior - -The launcher writes one JSON report and exits **0 only if all selected configurations pass**. -It exits **1** for any invalid selection, catalog, unreadable profile, or invalid profile. -An empty batch or an invalid/unreadable/ambiguous catalog fails the run with `Errors` populated -and no `Results`; no provider can be resolved safely in those cases. - -A successful two-provider report has this shape: - -```json -{ - "Valid": true, - "Errors": [], - "Results": [ - { "Index": 1, "Provider": "telesign", "Valid": true, "Code": "ConfigurationValid", "Issues": [] }, - { "Index": 2, "Provider": "soprano", "Valid": true, "Code": "ConfigurationValid", "Issues": [] } - ] -} -``` - -Results retain input order and a one-based `Index`. Unknown or empty provider selections have -`Provider: null` rather than echoing arbitrary input. Other result codes are `InvalidSelection`, -`DuplicateSelection`, `UnknownProvider`, `ProfileUnreadable`, and `ProfileInvalid`. `Issues` -contains safe descriptions; malformed input is not copied into an error message. - -For example, `-Provider telesign,sinch,soprano` returns three results: Telesign and Soprano can -pass independently, Sinch returns `UnknownProvider`, the overall `Valid` is `false`, and the -exit code is 1. Similarly, an invalid Telesign profile does not prevent Soprano from being checked. -Fix the reported profile fields or JSON syntax locally, then rerun the same command. - -For automation already running inside PowerShell, use the exported function without exiting -the caller; check `Valid` explicitly: - -```powershell -Import-Module .\setup\support\Epp.Setup.psm1 -$report = Test-EppProviderConfiguration -Provider telesign,soprano -Channel sms -EndpointRegion global -if (-not $report.Valid) { - $report | ConvertTo-Json -Depth 6 - throw 'Offline provider configuration validation failed.' -} -``` - -## Code and validation boundary - -- [CLI entry point](../setup/Test-EppProviders.ps1): JSON serialization and exit status. -- [Shared setup module](../setup/support/Epp.Setup.psm1): `Test-EppProviderConfiguration` resolves - independent selections using `Get-EppProviderCatalog`, then reuses - `ConvertTo-EppProviderSettings`, the same validator used by single-provider setup. -- [Focused tests](../setup/tests/Providers.Tests.ps1): independent results, invalid/duplicate - selections, safe errors, offline behavior, CLI status, and single-provider regression. - -The shared validator checks the enabled flag, provider identity, tenant GUID, authentication mode, -Key Vault **secret-name syntax** for API-key profiles, public HTTPS endpoint syntax, timeout/retry -bounds, and OAuth application ID/scope syntax. It requires **all four SMS/voice Global/EU routes**, -including routes not selected for this run, just as deployment setup does. Public hostname syntax -is not a DNS or reachability test. Secret names are configuration references, not fetched secrets. - -This command only reads local catalog/profile files. It performs no HTTP requests, provider -sends, credential reads/refresh/rotation, resource deployments, or Graph policy/role changes. -It does not contact SAS, call the Function, or rely on evaluation/fallback behavior. - -Existing deployment remains **one provider per endpoint**, selected by `EPP_PROVIDER_NAME`. -No runtime switching, fan-out, fallback, or multi-provider deployment is introduced. The -authenticated encrypted evaluation contract is unchanged: valid `mode: 2` requests return the -nonce before provider selection, credential resolution, or outbound provider calls. Offline -validation does not exercise or prove that contract on a deployed endpoint. Background credential -refresh in an existing deployment remains independent and is not triggered by this command. - -Deployment, authenticated evaluation, provider onboarding, and any separately authorized live -delivery checks remain distinct steps. Never interpret `ConfigurationValid` as permission to -send a message or as evidence that an exposed credential is safe to reuse. diff --git a/setup/Test-EppProviders.ps1 b/setup/Test-EppProviders.ps1 deleted file mode 100644 index ab883f0..0000000 --- a/setup/Test-EppProviders.ps1 +++ /dev/null @@ -1,31 +0,0 @@ -#Requires -Version 7.0 -<# -.SYNOPSIS - Validate multiple local setup provider profiles without network or credential access. -.DESCRIPTION - Writes one JSON report. Exits 0 only when every selected provider configuration is valid; - otherwise exits 1. This does not test provider credentials, entitlement, or delivery. -.EXAMPLE - .\Test-EppProviders.ps1 -Provider telesign,soprano -Channel sms -EndpointRegion global -#> -[CmdletBinding()] -param( - [AllowEmptyCollection()] [AllowEmptyString()] [string[]] $Provider = @(), - [string] $Channel, - [string] $EndpointRegion, - [string] $ProfileDirectory = (Join-Path $PSScriptRoot 'providers') -) - -$ErrorActionPreference = 'Stop' -Set-StrictMode -Version Latest -$module = Import-Module (Join-Path $PSScriptRoot 'support/Epp.Setup.psm1') -PassThru -Force -try { - $report = Test-EppProviderConfiguration -Provider $Provider -Channel $Channel ` - -EndpointRegion $EndpointRegion -ProfileDirectory $ProfileDirectory - $report | ConvertTo-Json -Depth 6 - if (-not $report.Valid) { exit 1 } -} -finally { - Remove-Module -ModuleInfo $module -} -exit 0 diff --git a/setup/support/Epp.Setup.psm1 b/setup/support/Epp.Setup.psm1 index 80e0ca6..981b847 100644 --- a/setup/support/Epp.Setup.psm1 +++ b/setup/support/Epp.Setup.psm1 @@ -133,37 +133,6 @@ function Read-EppInput { } } -function Get-EppProviderCatalog { - param([string] $ProfileDirectory) - - $catalog = Read-EppJson (Join-Path $ProfileDirectory 'catalog.json') - if ($catalog['schemaVersion'] -ne 1 -or -not $catalog['providers']) { throw 'Unsupported or empty provider catalog.' } - $entries = @($catalog['providers']) - $ids = @{} - $aliases = @{} - $files = @{} - foreach ($entry in $entries) { - if ($entry -isnot [Collections.IDictionary] -or - $entry['id'] -isnot [string] -or $entry['id'] -cnotmatch '^[a-z][a-z0-9-]{1,31}$' -or - $entry['file'] -isnot [string] -or - $entry['file'] -cnotmatch '^[a-z][a-z0-9-]{1,31}\.json$' -or - $entry['displayName'] -isnot [string] -or [string]::IsNullOrWhiteSpace($entry['displayName']) -or - $entry['displayName'] -match '[\x00-\x1f]' -or $ids.ContainsKey($entry['id']) -or - $files.ContainsKey($entry['file'])) { - throw 'Provider catalog contains an invalid or duplicate entry.' - } - $ids[$entry['id']] = $true - $files[$entry['file']] = $true - foreach ($alias in @($entry['id'], $entry['displayName'])) { - if ($aliases.ContainsKey($alias) -and $aliases[$alias] -ne $entry['id']) { - throw 'Provider catalog contains an ambiguous ID or display name.' - } - $aliases[$alias] = $entry['id'] - } - } - return $entries -} - function Get-EppProvider { param( [string] $AssetDirectory, [string] $SourceBaseUri, [string] $Provider, [string] $Channel, @@ -176,7 +145,18 @@ function Get-EppProvider { if ($SourceBaseUri -cnotmatch $sourcePattern) { throw 'Provider files must come from the same commit-pinned selected repository as the deployment tools.' } - $entries = @(Get-EppProviderCatalog -ProfileDirectory (Join-Path $AssetDirectory 'providers')) + $catalog = Read-EppJson (Join-Path $AssetDirectory 'providers/catalog.json') + if ($catalog['schemaVersion'] -ne 1 -or -not $catalog['providers']) { throw 'Unsupported or empty provider catalog.' } + $entries = @($catalog['providers']) + $ids = @{} + foreach ($entry in $entries) { + if ($entry -isnot [Collections.IDictionary] -or $entry['id'] -cnotmatch '^[a-z][a-z0-9-]{1,31}$' -or + $entry['file'] -cnotmatch '^[a-z][a-z0-9-]{1,31}\.json$' -or + -not $entry['displayName'] -or $entry['displayName'] -match '[\x00-\x1f]' -or $ids.ContainsKey($entry['id'])) { + throw 'Provider catalog contains an invalid or duplicate entry.' + } + $ids[$entry['id']] = $true + } $selected = Select-EppOption -Entries $entries -Name Provider -Value $Provider -NonInteractive:$NonInteractive $path = Join-Path $AssetDirectory "providers/$($selected['file'])" @@ -256,9 +236,7 @@ function ConvertTo-EppProviderSettings { } } if ($issues.Count) { - $validationError = [IO.InvalidDataException]::new("Provider '$DisplayName' is not deployment-ready:`n - $($issues -join "`n - ")`nAsk the provider owner to complete its GitHub JSON. No Azure resources were changed.") - $validationError.Data['EppValidationIssues'] = $issues.ToArray() - throw $validationError + throw "Provider '$DisplayName' is not deployment-ready:`n - $($issues -join "`n - ")`nAsk the provider owner to complete its GitHub JSON. No Azure resources were changed." } $channelEntry = Select-EppOption -Entries @( @@ -295,85 +273,6 @@ function ConvertTo-EppProviderSettings { } } -function Test-EppProviderConfiguration { - [CmdletBinding()] - param( - [AllowEmptyCollection()] [AllowEmptyString()] [string[]] $Provider = @(), - [string] $Channel, - [string] $EndpointRegion, - [string] $ProfileDirectory = (Join-Path $PSScriptRoot '../providers') - ) - - $report = [pscustomobject]@{ Valid = $false; Errors = @(); Results = @() } - if ($null -eq $Provider -or -not $Provider.Count) { - $report.Errors = @('Select at least one provider catalog ID.') - return $report - } - try { $entries = @(Get-EppProviderCatalog -ProfileDirectory $ProfileDirectory) } - catch { - $report.Errors = @('Cannot read a valid, unambiguous local provider catalog. Check catalog.json and its entries.') - return $report - } - - $selections = @($Provider | ForEach-Object { ([string]$_).Trim().ToLowerInvariant() }) - $results = for ($index = 0; $index -lt $selections.Count; $index++) { - $id = $selections[$index] - $selected = $entries | Where-Object { $_['id'] -ceq $id } - $result = [pscustomobject]@{ - Index = $index + 1 - Provider = if ($selected) { $selected['id'] } else { $null } - Valid = $false - Code = 'InvalidSelection' - Issues = @() - } - if (-not $id) { - $result.Issues = @('Provider selection must not be empty.') - } - elseif (@($selections | Where-Object { $_ -ceq $id }).Count -gt 1) { - $result.Code = 'DuplicateSelection' - $result.Issues = @('Each provider catalog ID may be selected only once (case-insensitive).') - } - elseif (-not $selected) { - $result.Code = 'UnknownProvider' - $result.Issues = @('Provider selection must match an ID in the local setup catalog; aliases and runtime-only adapters are not supported.') - } - elseif ($Channel -notin @('sms', 'voice') -or $EndpointRegion -notin @('global', 'eu')) { - $result.Issues = @('Specify Channel sms or voice and EndpointRegion global or eu.') - } - else { - try { - $profile = Read-EppJson (Join-Path $ProfileDirectory $selected['file']) - } - catch { - $result.Code = 'ProfileUnreadable' - $result.Issues = @('Cannot read the local profile as a JSON object. Check the catalog file entry and JSON syntax.') - $result - continue - } - try { - $null = ConvertTo-EppProviderSettings -Profile $profile -Id $selected['id'] -DisplayName $selected['id'] ` - -Channel $Channel -EndpointRegion $EndpointRegion -NonInteractive - $result.Valid = $true - $result.Code = 'ConfigurationValid' - } - catch { - $result.Code = 'ProfileInvalid' - # Only validator-authored field descriptions are safe to return, not parser errors or input values. - $result.Issues = @( - if ($_.Exception.Data.Contains('EppValidationIssues')) { - $_.Exception.Data['EppValidationIssues'] - } - else { 'Profile structure is invalid. Check deployment, authentication, and all SMS/voice Global/EU routes.' } - ) - } - } - $result - } - $report.Results = @($results) - $report.Valid = @($report.Results | Where-Object { -not $_.Valid }).Count -eq 0 - return $report -} - function Get-EppResourceNames { param([string] $SubscriptionId, [string] $ApplicationId, [string] $ResourcePrefix) @@ -1785,4 +1684,4 @@ function Invoke-EppSetup { } } -Export-ModuleMember -Function Invoke-EppSetup, Test-EppProviderConfiguration +Export-ModuleMember -Function Invoke-EppSetup diff --git a/setup/tests/Providers.Tests.ps1 b/setup/tests/Providers.Tests.ps1 deleted file mode 100644 index 1b0962f..0000000 --- a/setup/tests/Providers.Tests.ps1 +++ /dev/null @@ -1,196 +0,0 @@ -#Requires -Version 7.0 -$ErrorActionPreference = 'Stop' -Set-StrictMode -Version Latest -$module = Import-Module (Join-Path $PSScriptRoot '../support/Epp.Setup.psm1') -Force -PassThru -$directory = Join-Path ([IO.Path]::GetTempPath()) "epp-provider-tests-$([Guid]::NewGuid().ToString('N'))" -$profiles = Join-Path $directory 'providers' -New-Item -ItemType Directory -Path $profiles | Out-Null -Copy-Item (Join-Path $PSScriptRoot '../providers/*.json') -Destination $profiles -try { - & $module { - param($Directory, $Profiles) - - function script:Assert($Condition, [string] $Message) { - if (-not $Condition) { throw $Message } - } - $script:ForbiddenCalls = [Collections.Generic.List[string]]::new() - foreach ($command in @('Invoke-WebRequest', 'Invoke-RestMethod', 'Invoke-EppAz', 'Invoke-MgGraphRequest', - 'Get-AzKeyVaultSecret', 'Get-Secret', 'Read-Host')) { - Set-Item -Path "Function:script:$command" -Value { - $script:ForbiddenCalls.Add($MyInvocation.MyCommand.Name) - throw 'Network, credentials and prompts are forbidden in offline validation.' - } - } - $script:AllowedReads = @('catalog.json', 'telesign.json', 'soprano.json') | - ForEach-Object { Join-Path $Profiles $_ } - function script:Get-Content { - param($LiteralPath, [switch] $Raw) - Assert ($LiteralPath -in $script:AllowedReads) 'Validation must only read local catalog/profile JSON, never credentials.' - Microsoft.PowerShell.Management\Get-Content -LiteralPath $LiteralPath -Raw:$Raw - } - $parameters = @{ Provider = @('telesign', 'soprano'); Channel = 'sms'; EndpointRegion = 'global'; ProfileDirectory = $Profiles } - foreach ($channel in @('sms', 'voice')) { - foreach ($region in @('global', 'eu')) { - $report = Test-EppProviderConfiguration @parameters -Channel $channel -EndpointRegion $region - Assert ($report.Valid -and $report.Results.Count -eq 2 -and $report.Errors.Count -eq 0) 'Both shipped profiles must pass for every channel/region.' - Assert (($report.Results.Provider -join ',') -ceq 'telesign,soprano') 'Results must retain selection order and canonical provider IDs.' - Assert (@($report.Results | Where-Object { $_.Code -ne 'ConfigurationValid' -or $_.Issues.Count }).Count -eq 0) 'Successful results must contain no issues.' - } - } - $single = Test-EppProviderConfiguration @parameters -Provider ' TELESIGN ' - Assert ($single.Valid -and $single.Results.Count -eq 1 -and $single.Results[0].Provider -ceq 'telesign') 'Single selections must normalize case and surrounding whitespace.' - - foreach ($selection in @(@(), @(''), @(' '))) { - $report = Test-EppProviderConfiguration @parameters -Provider $selection - Assert (-not $report.Valid) 'Empty selection must fail without prompting or choosing a provider.' - } - $report = Test-EppProviderConfiguration @parameters -Provider $null - Assert (-not $report.Valid -and $report.Errors.Count -eq 1) 'Null selection must be an explicit failed report.' - $report = Test-EppProviderConfiguration @parameters -Provider @('telesign', '') - Assert ($report.Results[0].Valid -and $report.Results[1].Code -eq 'InvalidSelection' -and -not $report.Valid) 'An empty entry must not prevent another provider from being validated.' - $report = Test-EppProviderConfiguration @parameters -Provider @('telesign', ' TELESIGN ', 'soprano') - Assert (-not $report.Valid -and $report.Results.Count -eq 3) 'Duplicate selections must fail the run.' - Assert ($report.Results[0].Code -eq 'DuplicateSelection' -and $report.Results[1].Code -eq 'DuplicateSelection' -and $report.Results[2].Valid) 'Every occurrence of a duplicate must fail, without skipping independent providers.' - $report = Test-EppProviderConfiguration @parameters -Provider @('infobip', 'sinch', 'soprano', 'not-a-provider') - Assert (-not $report.Valid -and $report.Results[2].Valid) 'Runtime-only and unknown providers must fail without stopping known providers.' - Assert (@($report.Results | Where-Object Code -eq 'UnknownProvider').Count -eq 3) 'Unknown selections must be identified explicitly.' - foreach ($invalidRoute in @(@{ Channel = '' }, @{ Channel = 'fax' }, @{ EndpointRegion = '' }, @{ EndpointRegion = 'westus2' })) { - $arguments = $parameters.Clone() - foreach ($key in $invalidRoute.Keys) { $arguments[$key] = $invalidRoute[$key] } - $report = Test-EppProviderConfiguration @arguments - Assert (-not $report.Valid -and @($report.Results | Where-Object Code -eq 'InvalidSelection').Count -eq 2) 'Invalid/missing routes must not silently select a default.' - } - - $telesignPath = Join-Path $Profiles 'telesign.json' - $original = Get-Content -LiteralPath $telesignPath -Raw - $catalogPath = Join-Path $Profiles 'catalog.json' - $catalogOriginal = Get-Content -LiteralPath $catalogPath -Raw - $sentinel = 'DO-NOT-REPORT-PRIVATE-INPUT' - $profile = $original | ConvertFrom-Json -AsHashtable - $profile.deployment.routes.voice.eu.endpoint = "https://user:$sentinel@provider.contoso.com/?token=$sentinel" - $profile.deployment.routes.sms.global.timeoutMilliseconds = 2501 - $profile.deployment.routes.sms.global.retryIntervalSeconds = -1 - $profile.deployment.authentication.keyVaultSecretName = $sentinel - $profile | ConvertTo-Json -Depth 10 | Set-Content -LiteralPath $telesignPath - $report = Test-EppProviderConfiguration @parameters - Assert (-not $report.Valid -and -not $report.Results[0].Valid -and $report.Results[1].Valid) 'A bad profile must not stop validation of a good profile.' - Assert ($report.Results[0].Code -eq 'ProfileInvalid') 'Invalid profiles must be reported explicitly.' - $json = $report | ConvertTo-Json -Depth 6 - Assert (-not $json.Contains($sentinel)) 'Reports must never echo invalid profile values.' - Assert ($json.Contains('deployment.routes.voice.eu.endpoint') -and $json.Contains('timeoutMilliseconds') -and $json.Contains('retryIntervalSeconds') -and $json.Contains('keyVaultSecretName')) 'Safe issue descriptions must identify invalid fields, including unselected routes.' - - foreach ($badJson in @("{ malformed-$sentinel", 'null', '[]')) { - Set-Content -LiteralPath $telesignPath -Value $badJson - $report = Test-EppProviderConfiguration @parameters - Assert (-not $report.Valid -and $report.Results[0].Code -eq 'ProfileUnreadable' -and $report.Results[1].Valid) 'Malformed/nonobject JSON must fail independently.' - Assert (-not ($report | ConvertTo-Json -Depth 6).Contains($sentinel)) 'Parser exception details must not leak input values.' - } - foreach ($badProfile in @('{}', '{"deployment":{}}', '{"deployment":{"routes":[]}}')) { - Set-Content -LiteralPath $telesignPath -Value $badProfile - $report = Test-EppProviderConfiguration @parameters - Assert ($report.Results[0].Code -eq 'ProfileInvalid' -and $report.Results[1].Valid) 'Missing profile structure must be a safe independent failure.' - Assert ($report.Results[0].Issues -is [array]) 'Even a single structural issue must serialize as a JSON array.' - } - Remove-Item -LiteralPath $telesignPath - $report = Test-EppProviderConfiguration @parameters - Assert ($report.Results[0].Code -eq 'ProfileUnreadable' -and $report.Results[1].Valid) 'A missing profile must not trigger a download.' - Set-Content -LiteralPath $telesignPath -Value $original - foreach ($mutation in @('disabled', 'identity', 'tenant', 'authentication', 'timeout', 'retry')) { - $profile = $original | ConvertFrom-Json -AsHashtable - switch ($mutation) { - 'disabled' { $profile.deployment.enabled = $false } - 'identity' { $profile.deployment.providerName = 'soprano' } - 'tenant' { $profile.deployment.tenantId = $sentinel } - 'authentication' { $profile.deployment.authentication.mode = $sentinel } - 'timeout' { $profile.deployment.routes.sms.global.timeoutMilliseconds = 0 } - 'retry' { $profile.deployment.routes.sms.global.retryIntervalSeconds = 2147484 } - } - $profile | ConvertTo-Json -Depth 10 | Set-Content -LiteralPath $telesignPath - $report = Test-EppProviderConfiguration @parameters - Assert ($report.Results[0].Code -eq 'ProfileInvalid' -and $report.Results[1].Valid) "Shared validator must reject $mutation independently." - Assert ($report.Results[0].Issues -is [array] -and $report.Results[0].Issues.Count -eq 1) 'Single field issues must retain array shape.' - Assert (-not ($report | ConvertTo-Json -Depth 6).Contains($sentinel)) 'Invalid tenant/authentication values must not leak.' - } - Set-Content -LiteralPath $telesignPath -Value $original - $sopranoPath = Join-Path $Profiles 'soprano.json' - $sopranoOriginal = Get-Content -LiteralPath $sopranoPath -Raw - $profile = $sopranoOriginal | ConvertFrom-Json -AsHashtable - $profile.deployment.routes.voice.eu.appId = $sentinel - $profile.deployment.routes.voice.eu.scope = $sentinel - $profile | ConvertTo-Json -Depth 10 | Set-Content -LiteralPath $sopranoPath - $report = Test-EppProviderConfiguration @parameters - Assert (-not $report.Valid -and $report.Results[0].Valid -and $report.Results[1].Code -eq 'ProfileInvalid') 'Invalid OAuth settings must fail independently, including unselected routes.' - Assert ($report.Results[1].Issues.Count -eq 2 -and -not ($report | ConvertTo-Json -Depth 6).Contains($sentinel)) 'OAuth issues must use safe field descriptions.' - Set-Content -LiteralPath $sopranoPath -Value $sopranoOriginal - - foreach ($mutation in @('duplicate-id', 'duplicate-file', 'ambiguous-alias', 'path-traversal', 'schema')) { - $catalog = $catalogOriginal | ConvertFrom-Json -AsHashtable - switch ($mutation) { - 'duplicate-id' { $catalog.providers[1].id = 'telesign' } - 'duplicate-file' { $catalog.providers[1].file = 'telesign.json' } - 'ambiguous-alias' { $catalog.providers[1].displayName = 'Telesign' } - 'path-traversal' { $catalog.providers[0].file = '../private.json' } - 'schema' { $catalog.schemaVersion = 2 } - } - $catalog | ConvertTo-Json -Depth 6 | Set-Content -LiteralPath $catalogPath - $report = Test-EppProviderConfiguration @parameters - Assert (-not $report.Valid -and $report.Results.Count -eq 0 -and $report.Errors.Count -eq 1) 'Invalid or ambiguous catalogs must fail the entire run before profile reads.' - } - Set-Content -LiteralPath $catalogPath -Value "{ malformed-$sentinel" - $report = Test-EppProviderConfiguration @parameters - Assert (-not $report.Valid -and -not ($report | ConvertTo-Json -Depth 6).Contains($sentinel)) 'Catalog parser failures must not leak input.' - Set-Content -LiteralPath $catalogPath -Value $catalogOriginal - $report = Test-EppProviderConfiguration @parameters -Provider @($sentinel, 'soprano') - Assert (-not $report.Valid -and $report.Results[0].Index -eq 1 -and $null -eq $report.Results[0].Provider -and $report.Results[1].Valid) 'Unknown selections must be identified by position rather than echoed.' - Assert (-not ($report | ConvertTo-Json -Depth 6).Contains($sentinel)) 'Unknown selections must not leak arbitrary input.' - Assert ($script:ForbiddenCalls.Count -eq 0) 'Offline validation must not attempt network, secret access, or prompts even on failures.' - - $script:Downloads = @() - function script:Invoke-WebRequest { - param($Uri, $OutFile, $TimeoutSec, $MaximumRedirection) - $script:Downloads += $Uri - } - foreach ($id in @('telesign', 'soprano')) { - $selected = Get-EppProvider -AssetDirectory $Directory ` - -SourceBaseUri "https://raw.githubusercontent.com/Azure-Samples/ExternalPhoneProvider-AzureFunction-Sample/$('a' * 40)/setup" ` - -Provider $id -Channel voice -EndpointRegion eu -NonInteractive - $expected = ConvertTo-EppProviderSettings -Profile (Read-EppJson (Join-Path $Profiles "$id.json")) ` - -Id $id -DisplayName $id -Channel voice -EndpointRegion eu -NonInteractive - foreach ($key in $expected.Settings.Keys) { - Assert ($selected.Settings[$key] -ceq $expected.Settings[$key]) 'Single-provider setup must preserve its existing app settings.' - } - Assert ($selected.Settings.EPP_PROVIDER_NAME -ceq $id -and $selected.Channel -eq 'voice' -and $selected.EndpointRegion -eq 'eu') 'Setup must still select exactly one provider and route.' - } - Assert ($script:Downloads.Count -eq 2) 'Existing setup must retain one pinned profile download per selection.' - } $directory $profiles - - $launcher = (Join-Path $PSScriptRoot '../Test-EppProviders.ps1').Replace("'", "''") - $escapedProfiles = $profiles.Replace("'", "''") - foreach ($case in @( - @{ Selection = 'telesign,soprano'; Exit = 0; Count = 2 }, - @{ Selection = 'telesign'; Exit = 0; Count = 1 }, - @{ Selection = 'telesign,unknown'; Exit = 1; Count = 2 }, - @{ Selection = 'telesign,TELESIGN'; Exit = 1; Count = 2 }, - @{ Selection = '@()'; Exit = 1; Count = 0 }, - @{ Selection = 'telesign,soprano'; Exit = 1; Count = 2; InvalidProfile = $true } - )) { - if ($case.ContainsKey('InvalidProfile')) { - $path = Join-Path $profiles 'telesign.json' - $profile = Get-Content -LiteralPath $path -Raw | ConvertFrom-Json -AsHashtable - $profile.deployment.enabled = $false - $profile | ConvertTo-Json -Depth 10 | Set-Content -LiteralPath $path - } - $output = & (Join-Path $PSHOME 'pwsh') -NoProfile -Command "& '$launcher' -Provider $($case.Selection) -Channel sms -EndpointRegion global -ProfileDirectory '$escapedProfiles'" 2>&1 - if ($LASTEXITCODE -ne $case.Exit) { throw "Unexpected CLI exit code for $($case.Selection): $LASTEXITCODE" } - $report = ($output -join "`n") | ConvertFrom-Json - if ($report.Valid -ne ($case.Exit -eq 0) -or $report.Results.Count -ne $case.Count) { throw 'CLI must emit exactly one JSON report matching its exit code.' } - foreach ($result in $report.Results) { - if ($result.Issues -isnot [array]) { throw 'CLI issues must always be JSON arrays, including single-issue failures.' } - } - } - Write-Host 'Provider validation tests passed (offline batch, safe failures, CLI exit codes, single-provider regression).' -} -finally { - Remove-Module -ModuleInfo $module - Remove-Item -LiteralPath $directory -Recurse -Force -} From 64c03fa3f1abb92b2d774a3f7d7787ad897c52a1 Mon Sep 17 00:00:00 2001 From: Hou Chi Chan Date: Thu, 8 Oct 2026 10:54:34 -0700 Subject: [PATCH 03/14] Make multi-provider guide a step-by-step walkthrough Lead with practical JavaScript integration steps, small fenced examples and source links. Keep account isolation and live-test boundaries clear while shortening the experiment narrative. Co-authored-by: Copilot App <223556219+Copilot@users.noreply.github.com> Copilot-Session: 9314a109-37df-4021-8e75-bea995d57b23 --- docs/MULTI-PROVIDER-IMPLEMENTATION.md | 532 ++++++++++++++------------ 1 file changed, 295 insertions(+), 237 deletions(-) diff --git a/docs/MULTI-PROVIDER-IMPLEMENTATION.md b/docs/MULTI-PROVIDER-IMPLEMENTATION.md index b5e31c3..fb82b2a 100644 --- a/docs/MULTI-PROVIDER-IMPLEMENTATION.md +++ b/docs/MULTI-PROVIDER-IMPLEMENTATION.md @@ -1,128 +1,50 @@ -# Implementing multiple provider configurations - -**A customer customization is feasible, but it is not a built-in feature.** An offline JavaScript -prototype exercised the existing adapters with independently configured Telesign API-key and -Soprano OAuth contexts, including concurrent requests. The critical requirement is to isolate -configuration and credentials per context, not merely change the selected provider name. - -This guide describes changes **you would implement in your own application**. It does not ship a -router, setup script, configuration validator, or new deployment behavior. The unmodified sample -still uses **one provider per deployment**, selected by `EPP_PROVIDER_NAME`. Separate deployments -remain the simpler option when accounts require strong isolation. - -## What was demonstrated offline - -A session-only experiment against the JavaScript implementation at commit -`0d5db07f00a8b83295d83344a877f3bfd1ef009a`, using Node.js 22.17.1, passed six focused tests: - -- Forty concurrent synthetic SMS dispatches used four named contexts: two Telesign accounts and - two Soprano accounts. Each had its own frozen configuration and `CredentialTokenService`. - Assertions checked the exact endpoint, account/token, request content, correlation ID and - response mapping for every dispatch. Concurrent credential acquisition coalesced within each - context, not across accounts. -- A negative control deliberately reused one credential service with another account and then - another authentication mode. It returned the first cached API-key bundle. **Changing the - arguments passed to the existing singleton does not switch its configuration.** -- A locally encrypted evaluation request was parsed and decrypted, then returned its nonce - before route lookup, credential acquisition or transport, even with an unknown route ID. - An incomplete decrypted context still failed before routing. -- Unknown/duplicate route IDs, an unknown adapter and a channel mismatch were rejected before - credential or transport calls. -- A synthetic provider HTTP failure was terminal for that request while an independent - concurrent route succeeded; no retry or alternate-provider dispatch occurred. -- After simulated secret-cache expiry, a failing account could not borrow another context's - credentials or cause fallback. An independent OAuth route still succeeded. - -The prototype used the **real** configuration reader, payload parser, local JWE decryption, -credential service/cache classes, Telesign/Soprano request builders, shared transport logic and -response mappers. Only the Azure SDK boundary and `fetch` were replaced with deterministic fakes; -socket, HTTP and DNS access were guarded to fail. Periodic refresh scheduling was simulated and -every service was closed. No process environment was changed per request. Existing JavaScript -adapter, credential-cache and handler tests also passed. - -**This was not an end-to-end provider or SAS test.** The prototype used a customer-style dispatch -seam, test-owned route IDs, fake credentials, fake responses and a locally generated encryption key. -It did not run a deployed Function, validate Easy Auth, establish a trusted routing identity, -perform a real OAuth exchange, read Key Vault, or send SMS/voice messages. Its new multi-context -dispatch tests covered SMS, not a multi-provider voice rollout. Real refresh scheduling, distributed -workers, production load and configuration reload were not demonstrated. The prototype is not -included as a supported repository command; the evidence does not certify a production router. - -## Start at these existing code seams - -### JavaScript: experimentally exercised - -- [SendOtp.js](../javascript/src/functions/SendOtp.js): the registered handler validates/decrypts, - returns evaluation before selection, checks channel/authentication/endpoint, resolves credentials, - builds one request, calls transport and maps the endpoint response. Its - `startProviderCredentialRefresh`/`stopProviderCredentialRefresh` hooks currently manage one - selected credential service. -- [config.js](../javascript/src/functions/config.js): `readConfig(env)` / `AppConfig` already accept - an explicit settings object; use that seam rather than mutating `process.env`. -- [providers/index.js](../javascript/src/functions/providers/index.js): `selectProvider(name)` - resolves a fixed adapter with no default. [Telesign](../javascript/src/functions/providers/telesign.js) - and [Soprano](../javascript/src/functions/providers/soprano.js) expose `credentialSpec`, - `createRequest` and `interpretResponse`. -- [credentials.js](../javascript/src/functions/credentials.js): `CredentialTokenService`, - `ApiKeyCache` and `AccessTokenCache` own acquisition, cache expiry, in-flight refresh and shutdown. - The exported singleton owns **one** selected cache; it is not keyed by provider or configuration. -- [providerTransport.js](../javascript/src/functions/providerTransport.js): `sendProviderRequest` - enforces request URL checks, timeout and manual redirects. - [providerResult.js](../javascript/src/functions/providerResult.js) maps adapter outcomes to - endpoint statuses; preserve those mappings rather than treating all HTTP responses as success. -- [entraPayload.js](../javascript/src/functions/entraPayload.js), - [jwe.js](../javascript/src/functions/jwe.js) and - [logging.js](../javascript/src/functions/logging.js): retain payload/decryption checks and - selected-metadata logging. Parsing a `tenantId` field does not establish its routing authority. - -### .NET: inspected, not multi-context-tested here - -Follow [Functions/SendOtp.cs](../dotnet/Functions/SendOtp.cs) for evaluation, `SelectProvider` and -dispatch; [AppConfig.Read(IEnv)](../dotnet/Src/AppConfig.cs) and -[IEnv](../dotnet/Src/Models.cs) provide explicit configuration seams. -[PhoneProviderBase](../dotnet/Src/PhoneProviderBase.cs) defines credential acquisition, -`SendOtpAsync`, shared HTTP handling and response-status mapping. - -[CredentialTokenService](../dotnet/Src/CredentialTokenService.cs) caches by `provider.Name`, which -does not distinguish two accounts of the same adapter. -[SopranoProvider](../dotnet/Src/Providers/SopranoProvider.cs) also retains the first initialized -OAuth identity/credential/scope. [SecretResolver](../dotnet/Src/SecretResolver.cs) retains a vault -client built from its `IEnv`. Merely changing the outer cache key is therefore insufficient. -Build a context-owned service/provider/resolver/configuration object graph, or redesign every -relevant cache key and initialization boundary. Review cold-acquisition concurrency explicitly; -do not assume its cache has the same single-flight behavior as JavaScript. -[Program.cs](../dotnet/Program.cs) currently registers singleton providers/services and a -non-redirecting HTTP client; customer DI/lifecycle changes must preserve transport safety. - -### Python: inspected, not multi-context-tested here - -In [function_app.py](../python/function_app.py), `_select_provider`, `_send_to_provider` and the -evaluation branch are the dispatch seams. `_send_to_provider` currently rereads process settings. -Pass the selected context explicitly instead. [read_config(env)](../python/src/config.py) already -accepts a mapping. [CredentialTokenService](../python/src/credentials.py) owns one selected cache; -construct one per context, with a context-specific [SecretResolver(env)](../python/src/secrets.py). -Do not share the module-level `_credentials` across configurations. Update the warmup thread and -shutdown lifecycle deliberately. -[PhoneProviderBase](../python/src/provider.py) provides `build_request`, `map_response`, -`send_otp`, transport handling and endpoint-status mapping. - -The runtime adapters include Infobip, Sinch, Soprano and Telesign, but this experiment exercised -only Telesign and Soprano. The [setup catalog](../setup/providers/catalog.json) provisions only -Telesign/Soprano profiles. Adapter availability, setup coverage and a provider's approved account -capabilities are different things; none establishes a new multi-provider deployment contract. - -## Design your customization - -### Define named contexts and a trusted, deterministic selection rule - -A context ID identifies a complete configuration, not just an adapter name. For example, -`telesign-primary` and `telesign-secondary` must remain distinct even though both use Telesign. -Create a finite, reviewed context list at startup. Reject duplicate IDs, unknown adapters, -ambiguous rules, missing credentials configuration, disallowed endpoints and mismatched -authentication/channel settings before enabling live use. - -The following is **illustrative customer-owned configuration, not a supported new file format**. -Values in angle brackets must be supplied and reviewed; never place secret values in it: +# Add multiple provider configurations to your application + +You can adapt this sample to choose between multiple provider accounts in one application. +The key is to give each account its own configuration and credential cache, then select exactly +one account for each request. + +The sample does **not** do this out of the box: it still uses one provider per deployment through +`EPP_PROVIDER_NAME`. This walkthrough shows where to make the changes in your own JavaScript +application. It does not add a router or setup script to the repository. If accounts need a strong +security boundary, keep them in separate deployments instead. + +The snippets are **integration examples, not a drop-in handler**. They fit into +[SendOtp.js](../javascript/src/functions/SendOtp.js) and use its existing variables and helpers. +Anything named `CUSTOMER_*` is a placeholder you must implement; it is not a repository API. + +## 1. Choose a routing rule you control + +Start with a simple, server-owned rule. For example, use one account for SMS and another for voice: + +```javascript +// Module scope in your customized SendOtp.js, not supplied by the caller. +const routeByChannel = Object.freeze({ + sms: 'telesign-primary', + voice: 'soprano-primary', +}); +``` + +Open [entraPayload.js](../javascript/src/functions/entraPayload.js) to see how the handler +validates the channel and exposes `payload.channelName`. Use that validated value with your +approved policy. Reject unsupported or ambiguous routes; do not silently choose a default. + +If you need two accounts for the same channel, first decide how your application will verify +the customer or endpoint identity that owns the request. The existing contract has no trusted +provider-selector field. An arbitrary header, `providerId`, correlation ID or unverified body +`tenantId` is **not** authority to select an account. The SAS caller identity may not distinguish +your customers either. Keep endpoint URLs, vaults, scopes and secret names under server control. + +## 2. Define one configuration per provider account + +Open [config.js](../javascript/src/functions/config.js). Its `readConfig(env)` function accepts +an explicit settings object, so you can reuse it without changing `process.env` per request. + +Give each account a unique context ID. For example, two Telesign accounts would need different +IDs even though both have `EPP_PROVIDER_NAME: "telesign"`. + +Use the following as a starting point for **your own configuration format**. Replace the +placeholders with approved values. The sample does not load this JSON automatically. ```json [ @@ -155,129 +77,265 @@ Values in angle brackets must be supplied and reviewed; never place secret value ] ``` -For one endpoint, an explicit server-owned `sms -> telesign-primary`, -`voice -> soprano-primary` rule is one possible deterministic policy. Validate the incoming channel -using the existing contract and authorize it against that policy. This is a design example; the -experiment used test-supplied route IDs for SMS across all four contexts. - -If different customers/accounts need the same channel, first define and verify the **trusted -endpoint/customer binding** from which routing can be authorized. Do not assume the SAS caller -identity distinguishes every customer. Do not use an arbitrary request header, a newly invented -`providerId`, correlation ID, or unverified body `tenantId` as authority. The existing wire contract -does not supply a new trusted provider-selector field. Never allow a caller to supply an endpoint, -vault, scope or secret name. Unknown or ambiguous routing must fail closed without a default. - -### Keep state and credentials isolated - -Build each context from a copied, immutable settings object. In JavaScript, `readConfig` retains -the supplied `env` reference, so freeze that object as well as the returned config. Separate -inbound authentication/decryption settings from outbound provider settings; do not switch the -inbound key because a provider route changed. - -Create **one long-lived credential service per context**, not per request. Bind it permanently to -that context's adapter credential specification, vault/secret names, vault-reading identity, -OAuth tenant/resource scope, outbound application and outbound managed identity. A provider name -alone is not a safe cache key. Share nothing that can retain another account's token or secret. -For Telesign, the existing adapter uses `telesign-api-key` and `telesign-customer-id`; distinct vaults -can isolate accounts with those same names. Using different names in one vault requires an explicit -per-context credential-spec customization, not mutation of the shared adapter object. - -Keep expiry, acquisition bounds, failure reporting and refresh coalescing. A failed/expired -context must fail its own request, never borrow credentials. Bound the context count, close every -service at shutdown, and define how background refresh is started. A request's evaluation branch -must not resolve credentials; independently running warmup/refresh is a separate lifecycle concern. -For configuration changes, prefer a controlled worker restart or a versioned replacement of the -whole context with safe draining/disposal, rather than modifying a live cache's inputs. - -These are logical isolation boundaries, not tenant security boundaries: one compromised worker -may access all identities assigned to it. Review least-privilege identity/vault access and provider -onboarding separately; retain separate deployments where stronger isolation is required. - -### Insert selection after evaluation, then dispatch once - -The sketch below uses existing JavaScript interfaces but is **pseudocode, not a complete handler**. -The uppercase helpers are customer responsibilities; authorization, error handling, lifecycle and -response/logging behavior must be integrated with the existing handler rather than bypassed. +Store **references and identifiers, not secret values**. The +[Telesign adapter](../javascript/src/functions/providers/telesign.js) requests the secret names +`telesign-api-key` and `telesign-customer-id`. Separate vaults let two accounts use those same +names. Different names in one vault require your own per-context credential specification; +do not mutate the shared adapter's `credentialSpec`. + +The [Soprano adapter](../javascript/src/functions/providers/soprano.js) uses OAuth. Keep its +provider tenant, resource scope, outbound application and managed identity together as one +configuration. These settings do not grant consent or provider entitlement. + +Keep inbound authentication and decryption configuration separate. Selecting an outbound provider +must not change the key used to decrypt the incoming request. + +## 3. Create isolated, long-lived credential contexts + +Open [credentials.js](../javascript/src/functions/credentials.js). The exported +`credentialTokenService` singleton keeps **one selected cache**. Passing it a different account +or authentication mode does not switch that cache. + +Use a new `CredentialTokenService` for each context instead. Build the contexts once per worker, +not once per request. Add the class import alongside the existing imports in your customized +`SendOtp.js`; keep the other credential helpers that its logging code uses. ```javascript -// Startup: customer validates unique IDs, rules, endpoints, identities and required settings. +const { readConfig } = require('./config'); +const { selectProvider } = require('./providers'); +const { CredentialTokenService } = require('./credentials'); + +// CUSTOMER_LOAD_AND_VALIDATE_CONFIG is your startup-only configuration loader. +// It must validate the complete list before any context is used. +const entries = CUSTOMER_LOAD_AND_VALIDATE_CONFIG(); const contexts = new Map(); -for (const entry of VALIDATED_CUSTOMER_CONTEXTS) { + +for (const entry of entries) { + if (!entry.id || contexts.has(entry.id)) { + throw new Error('Missing or duplicate provider context ID'); + } const settings = Object.freeze({ ...entry.settings }); const config = Object.freeze(readConfig(settings)); const provider = selectProvider(config.providerName); - REQUIRE_VALID_CONTEXT(entry.id, provider, config, contexts); + if (!provider || config.providerAuthMode !== provider.authenticationMode) { + throw new Error('Invalid provider context'); + } contexts.set(entry.id, Object.freeze({ - config, provider, credentials: new CredentialTokenService() + config, + provider, + credentials: new CredentialTokenService(), })); } +``` + +Your loader must also validate required credential settings, allowed channels, approved endpoint +allowlists, timeout values, a bounded context count, and that every routing rule identifies exactly +one matching context. The checks above are not a complete configuration validator. +[providers/index.js](../javascript/src/functions/providers/index.js) supplies the fixed adapter +lookup; unknown names return `null`. + +Freeze both the settings and config objects: `readConfig` retains the settings object as `env`. +Never reuse a context's service for a different vault, account, OAuth scope or identity. A provider +name alone is not enough to identify cached credentials. + +## 4. Select the context only after evaluation returns + +In [SendOtp.js](../javascript/src/functions/SendOtp.js), find `if (evaluation)`. Leave that branch, +the preceding envelope checks, [JWE decryption](../javascript/src/functions/jwe.js), and the +complete-delivery-context check in place. Keep Easy Auth enabled at the application boundary. + +The order must remain: + +```text +Authenticate -> validate envelope -> decrypt -> check delivery context + -> evaluation? Return the existing nonce response + -> otherwise select one provider context +``` + +Immediately after the existing evaluation early return, replace the single-provider lookup with +your route lookup. This fragment uses the handler's existing `fail`, `respond`, `payload`, +`correlationId` and `requestId` variables: -// Request: retain Easy Auth, envelope validation and JWE checks from SendOtp. -const { payload, deliveryContext } = VALIDATE_AND_DECRYPT_WITH_EXISTING_HANDLER(request); -if (payload.isEvaluation) { - return EXISTING_EVALUATION_RESPONSE(deliveryContext.nonce); +```javascript +const contextId = routeByChannel[payload.channelName]; +const selectedContext = contexts.get(contextId); +if (!selectedContext + || selectedContext.config.providerChannel !== payload.channelName) { + fail('provider_selection', 'unknown_provider', 400); + return respond(400, { + error: 'provider_delivery_failed', correlationId, requestId, + }); } -const id = AUTHORIZE_AND_SELECT_ONE_CONTEXT(serverOwnedPolicy, payload.channelName); -const context = contexts.get(id); -REQUIRE_KNOWN_CONTEXT_AND_MATCHING_CHANNEL(context, payload.channelName); -const { config, provider, credentials } = context; -const credential = await credentials.getCredentials(provider.credentialSpec, config); -REQUIRE_COMPLETE_CREDENTIAL_FOR_ADAPTER(credential, provider); -const delivery = BUILD_EXISTING_OTP_DELIVERY(deliveryContext, payload); -const outbound = provider.createRequest({ - channel: payload.channelName, endpoint: config.providerEndpoint, - delivery, credential, env: config.env +const { provider, config: providerConfig, credentials } = selectedContext; +``` + +For customer/account-based routing, this is where your **verified, authorized** customer binding +must select the context instead of `routeByChannel`. Evaluation must not depend on that selection +or acquire provider credentials. An invalid encrypted evaluation still fails the existing checks; +do not add a shortcut that returns a nonce before decryption. + +## 5. Use the selected context for one dispatch + +Continue in [SendOtp.js](../javascript/src/functions/SendOtp.js). Keep its channel, authentication +mode and endpoint checks, but read their outbound settings from `providerConfig`. Keep using the +original inbound config for decryption and encryption-key checks. + +Inside the existing credential-resolution `try` block, replace the singleton call with: + +```javascript +credential = await credentials.getCredentials( + provider.credentialSpec, + providerConfig, +); +``` + +Keep the existing completeness checks for API-key identity/secret and OAuth access token, and +the `credential_unavailable` error path. For OAuth logging, pass `providerConfig` to +`credentialContext` too. + +Keep the handler's [OtpDelivery](../javascript/src/functions/delivery.js) construction so phone, +message, locale and correlation handling stay unchanged. In the existing request-build and +transport `try` blocks, use the selected endpoint, adapter settings and timeout: + +```javascript +providerRequest = provider.createRequest({ + channel, + endpoint: providerConfig.providerEndpoint, + delivery, + credential, + env: providerConfig.env, }); -const response = await sendProviderRequest( - outbound, parseProviderTimeout(config.providerTimeoutMs), safeLogContext); -const result = provider.interpretResponse(response); -return EXISTING_ENDPOINT_RESPONSE_MAPPING(result, deliveryContext.nonce); -// Shutdown: close each context.credentials; never fall through to another context on failure. + +// Keep this in the existing transport try/catch, not the request-build block. +transportResponse = await sendProviderRequest( + providerRequest, + parseProviderTimeout(providerConfig.providerTimeoutMs), + logContext, +); +``` + +The imports and error handling for `sendProviderRequest` and `parseProviderTimeout` already exist +in `SendOtp.js`. See [providerTransport.js](../javascript/src/functions/providerTransport.js) +for URL checks, manual redirects and timeout behavior. Preserve those controls and your approved +endpoint allowlist. Do not forward the inbound authorization header. + +Keep `provider.interpretResponse(transportResponse)` and the existing +[response mapping](../javascript/src/functions/providerResult.js). Preserve success/block/failure +statuses, the success nonce and correlation ID, and the handler's error responses. A provider +timeout, rejection or unknown response must not trigger another account or provider. +**One request gets one dispatch: no fan-out, automatic fallback or resend.** + +Retain the [logging helpers](../javascript/src/functions/logging.js) and their fixed failure +classifications. If you add a context ID to telemetry, use an approved non-secret identifier. +Never log OTP/phone data, request bodies, credentials, raw provider descriptions or exceptions. + +## 6. Update startup, refresh and shutdown together + +Find `startProviderCredentialRefresh` and `stopProviderCredentialRefresh` in +[SendOtp.js](../javascript/src/functions/SendOtp.js). They currently use the singleton. Replace +that lifecycle wiring as part of your customization; leaving it behind could warm the wrong account. + +Decide explicitly whether to prewarm approved contexts at startup or acquire credentials on the +first live request. A service starts its own periodic refresh when first used. Prewarming performs +credential I/O, so do not do it in an offline test without fake SDK boundaries. Report acquisition +failures through the existing safe reporting mechanism; never borrow another context's credentials. + +Close every context in the existing termination hook: + +```javascript +function stopProviderCredentialRefresh() { + for (const { credentials } of contexts.values()) { + credentials.close(); + } +} +``` + +Keep cache expiry, refresh coalescing and acquisition bounds. Use a controlled worker restart when +changing configuration, or implement safe draining and replacement of whole contexts. Do not +modify a live service's inputs. Evaluation requests must not start credential acquisition, but +already-running background refresh is independent of a request's evaluation branch. + +Separate contexts provide logical isolation, not a tenant security boundary. A compromised worker +may reach every identity assigned to it. Review least privilege and use separate deployments when +that risk is unacceptable. + +## 7. Test your changes offline first + +Start from the existing +[adapter tests](../javascript/test/provider-flow.test.js), +[credential-cache tests](../javascript/test/credential-cache.test.js) and +[handler tests](../javascript/test/sendotp.test.js). +They show how to fake credential acquisition and provider transport. Use synthetic delivery data +and a locally generated encryption key; make unexpected network calls fail. + +Run those existing tests from the repository root in PowerShell: + +```powershell +node --test .\javascript\test\provider-flow.test.js ` + .\javascript\test\credential-cache.test.js ` + .\javascript\test\sendotp.test.js ``` -Keep approved endpoint allowlists in addition to generic HTTPS checks, preserve redirect/timeout -controls, and never forward inbound authorization to a provider. Retain the existing success, -block, failure and nonce rules. Do not turn an unknown provider response into success. -Log an approved non-secret context identifier and fixed failure classification, not request bodies, -OTP/phone data, credentials, tokens, raw provider descriptions or exception contents. -This design chooses **one** context per request: no fan-out, automatic fallback or resend. - -## Validate your implementation and roll out separately - -1. Start with offline fixtures. Adapt the real interfaces in - [provider-flow.test.js](../javascript/test/provider-flow.test.js), - [credential-cache.test.js](../javascript/test/credential-cache.test.js) and - [sendotp.test.js](../javascript/test/sendotp.test.js). - Replace SDK acquisition and transport with fakes, guard network access, and use only synthetic - delivery data. Test both authentication modes and multiple accounts of the same adapter under - concurrency; assert endpoint, credentials and correlation isolation, not just success counts. -2. Add unknown/duplicate/ambiguous routes, invalid settings, channel mismatch, cross-account - attempts, failed/expired credential acquisition, provider rejection, timeout, malformed response, - shutdown and configuration replacement tests. Prove failures cause no second provider attempt. - Verify valid evaluation returns before routing/credentials and malformed evaluation still fails. - Extend voice and every additional adapter before claiming support. -3. Run your runtime's existing contracts as well. .NET examples are - [SendOtpTests](../dotnet/tests/SendOtpTests.cs), - [CredentialTokenServiceTests](../dotnet/tests/CredentialTokenServiceTests.cs) and - [ContractTests](../dotnet/tests/ContractTests.cs). Python examples are - [test_engine.py](../python/tests/test_engine.py), - [test_credential_cache.py](../python/tests/test_credential_cache.py), - [test_function_app.py](../python/tests/test_function_app.py) and - [test_contract.py](../python/tests/test_contract.py). The shared - [contract fixtures](../tests/fixtures/contract.json) help preserve response behavior. -4. Review identity permissions, provider entitlement/consent, route and sender approval, secret - lifecycle, authorization boundaries, rate limits, observability and operational ownership. - Resolve exposed-credential rotation before any live validation. Passing a local configuration - or mocked test establishes none of these. -5. Deploy only after separate approval to a dedicated nonproduction environment. Verify deployed - authentication, encryption, policy readback and authorized evaluation independently. A real SAS - trigger can have surrounding fallback behavior: do not assume evaluation is harmless merely - because this Function returns before its provider call. -6. Perform live provider/delivery validation only with explicit authorization, approved recipient, - coordinated policy/attempt limits and a one-attempt safety gate. Provider acceptance is not - proof of handset delivery. Use an explicit rollout/rollback decision; do not add automatic - fallback to conceal a failing integration. - -No Azure resources, Graph roles/policy, Key Vault credentials, SAS triggers or provider sends were -used or changed to produce this guide. The feasibility result concerns reusable code seams and -offline context isolation, not production readiness or a delivered multi-provider feature. +They are a starting point, not coverage for your new router. Add tests for concurrent requests to +different providers **and two accounts of the same provider**. Check the exact endpoint, +credential, message and correlation ID for each request. Test unknown/duplicate/ambiguous routes, +channel mismatch, expired or failing credentials, provider errors, timeouts, malformed responses, +shutdown and configuration replacement. Failures must cause no second dispatch. + +Check that valid evaluation returns before routing/credentials/transport and invalid evaluation +still fails. Cover voice and every adapter you intend to enable. The shared +[contract fixtures](../tests/fixtures/contract.json) help preserve endpoint response behavior. + +## 8. Review and approve live rollout separately + +Before deployment, review identity/vault permissions, OAuth consent, provider entitlement, +sender/channel approval, rate limits, secret rotation and operational ownership. Rotate any exposed +credentials first. Offline checks do not establish any of these prerequisites. + +Use a separately approved nonproduction deployment. Verify authentication, encryption and policy +readback before authorized evaluation. A real SAS trigger can involve surrounding fallback +behavior, so do not assume it is harmless because this Function's evaluation branch skips delivery. + +Only attempt live delivery with explicit authorization, an approved recipient, coordinated policy +and attempt limits, and a **one-attempt safety gate**. Provider acceptance is not proof of handset +delivery. Plan an explicit rollout and rollback; do not hide failures behind automatic fallback. +Nothing in this walkthrough authorizes a SAS trigger or provider send. + +## Using .NET or Python instead + +**.NET:** Start with [SendOtp.cs](../dotnet/Functions/SendOtp.cs) and +[AppConfig.Read(IEnv)](../dotnet/Src/AppConfig.cs). Isolate the entire configuration, provider, +credential service and secret resolver per account. +[CredentialTokenService](../dotnet/Src/CredentialTokenService.cs) caches by `provider.Name`; +[SopranoProvider](../dotnet/Src/Providers/SopranoProvider.cs) retains its initial OAuth +identity/scope, and [SecretResolver](../dotnet/Src/SecretResolver.cs) retains its initial vault +client. Changing only the outer cache key is not enough. Update the singleton registrations and +lifecycle in [Program.cs](../dotnet/Program.cs), preserve +[PhoneProviderBase](../dotnet/Src/PhoneProviderBase.cs) transport/response behavior, and test +concurrent cold acquisition rather than assuming JavaScript's coalescing behavior. +Extend [SendOtpTests](../dotnet/tests/SendOtpTests.cs) and +[CredentialTokenServiceTests](../dotnet/tests/CredentialTokenServiceTests.cs). + +**Python:** Start with `_send_to_provider` and the evaluation branch in +[function_app.py](../python/function_app.py). Pass a selected context explicitly instead of +rereading process settings. Use [read_config(env)](../python/src/config.py), a separate +[CredentialTokenService](../python/src/credentials.py) and +[SecretResolver(env)](../python/src/secrets.py) per account; the module-level credential service +owns one selected cache. Update warmup/shutdown and preserve +[PhoneProviderBase](../python/src/provider.py) transport/response behavior. Extend +[test_engine.py](../python/tests/test_engine.py), +[test_credential_cache.py](../python/tests/test_credential_cache.py) and +[test_function_app.py](../python/tests/test_function_app.py). + +## Scope of this guidance + +An offline JavaScript prototype checked isolated Telesign API-key and Soprano OAuth contexts, +concurrent SMS dispatch, failure isolation and evaluation ordering using fake SDK/transport +boundaries. It was not a deployed or authenticated end-to-end test. Multi-provider voice, +real OAuth/Key Vault, trusted customer routing, production refresh/scale-out and live delivery +still need your implementation and validation. .NET/Python pointers are based on code inspection, +not equivalent multi-context tests. + +Runtime adapters also include Infobip and Sinch; the [setup catalog](../setup/providers/catalog.json) +contains only Telesign and Soprano. Adapter availability does not imply setup coverage or account +entitlement. No Azure, Graph, Key Vault, SAS or provider operations were performed for this guide. From 5b4579f76622b66e2c29feb87a68088ca4d3446e Mon Sep 17 00:00:00 2001 From: Hou Chi Chan Date: Thu, 8 Oct 2026 10:55:56 -0700 Subject: [PATCH 04/14] Clarify provider configuration and lifecycle edits Keep inbound config untouched and show explicit context-based startup and shutdown replacements, including the lazy-acquisition alternative. Co-authored-by: Copilot App <223556219+Copilot@users.noreply.github.com> Copilot-Session: 9314a109-37df-4021-8e75-bea995d57b23 --- docs/MULTI-PROVIDER-IMPLEMENTATION.md | 29 +++++++++++++++++++++++++-- 1 file changed, 27 insertions(+), 2 deletions(-) diff --git a/docs/MULTI-PROVIDER-IMPLEMENTATION.md b/docs/MULTI-PROVIDER-IMPLEMENTATION.md index fb82b2a..0199a9a 100644 --- a/docs/MULTI-PROVIDER-IMPLEMENTATION.md +++ b/docs/MULTI-PROVIDER-IMPLEMENTATION.md @@ -178,7 +178,11 @@ do not add a shortcut that returns a nonce before decryption. Continue in [SendOtp.js](../javascript/src/functions/SendOtp.js). Keep its channel, authentication mode and endpoint checks, but read their outbound settings from `providerConfig`. Keep using the -original inbound config for decryption and encryption-key checks. +original `const config = readConfig()` for `config.decryptionKeyPem` and `config.expectedKeyId`. +Do not redeclare `config` in the handler or replace its inbound settings with the selected context. +After the evaluation return, change the existing checks' `config.providerChannel`, +`config.providerAuthMode` and `config.providerEndpoint` references to the corresponding +`providerConfig` properties; leave their failure branches intact. Inside the existing credential-resolution `try` block, replace the singleton call with: @@ -240,9 +244,23 @@ first live request. A service starts its own periodic refresh when first used. P credential I/O, so do not do it in an offline test without fake SDK boundaries. Report acquisition failures through the existing safe reporting mechanism; never borrow another context's credentials. -Close every context in the existing termination hook: +If you choose startup prewarming, replace **both existing function bodies** with the following +pattern. This assumes every context has passed your startup validation and is approved for +credential acquisition. Keep the existing `reportRefreshFailure` import; remove the singleton +import once no handler or lifecycle code references it. ```javascript +async function startProviderCredentialRefresh() { + for (const { provider, config: providerConfig, credentials } of contexts.values()) { + try { + await credentials.getCredentials(provider.credentialSpec, providerConfig); + } catch { + // Refresh failures are reported by the service; report initialization failures here. + if (!credentials.current) reportRefreshFailure('configuration'); + } + } +} + function stopProviderCredentialRefresh() { for (const { credentials } of contexts.values()) { credentials.close(); @@ -250,6 +268,13 @@ function stopProviderCredentialRefresh() { } ``` +Keep the existing `app.hook.appStart(startProviderCredentialRefresh)` and +`app.hook.appTerminate(stopProviderCredentialRefresh)` registrations once each; do not add duplicate +hooks. Build and validate `contexts` before startup runs. If you choose lazy acquisition instead, +remove the old app-start registration and its singleton warmup function, but retain the +context-closing termination hook above. In either case, no lifecycle code should continue to +resolve credentials from the original process-wide provider selection. + Keep cache expiry, refresh coalescing and acquisition bounds. Use a controlled worker restart when changing configuration, or implement safe draining and replacement of whole contexts. Do not modify a live service's inputs. Evaluation requests must not start credential acquisition, but From 5c4f61c13a760261e421fef3baf9595f900c3b0c Mon Sep 17 00:00:00 2001 From: Hou Chi Chan Date: Thu, 8 Oct 2026 10:57:03 -0700 Subject: [PATCH 05/14] Clarify existing imports in provider walkthrough Consolidate imports without duplicate declarations and retain credential reporting helpers during request and lifecycle migration. Co-authored-by: Copilot App <223556219+Copilot@users.noreply.github.com> Copilot-Session: 9314a109-37df-4021-8e75-bea995d57b23 --- docs/MULTI-PROVIDER-IMPLEMENTATION.md | 9 ++++++--- 1 file changed, 6 insertions(+), 3 deletions(-) diff --git a/docs/MULTI-PROVIDER-IMPLEMENTATION.md b/docs/MULTI-PROVIDER-IMPLEMENTATION.md index 0199a9a..c1c05b9 100644 --- a/docs/MULTI-PROVIDER-IMPLEMENTATION.md +++ b/docs/MULTI-PROVIDER-IMPLEMENTATION.md @@ -97,13 +97,16 @@ Open [credentials.js](../javascript/src/functions/credentials.js). The exported or authentication mode does not switch that cache. Use a new `CredentialTokenService` for each context instead. Build the contexts once per worker, -not once per request. Add the class import alongside the existing imports in your customized -`SendOtp.js`; keep the other credential helpers that its logging code uses. +not once per request. Consolidate the imports below with the existing imports in your customized +`SendOtp.js`: `readConfig` and `selectProvider` are already imported, so do not paste duplicate +`const` declarations. Add `CredentialTokenService` to the existing credentials import and retain +required logging helpers such as `reportRefreshFailure`. Remove the `credentialTokenService` +singleton import only after replacing all its request and lifecycle uses. ```javascript const { readConfig } = require('./config'); const { selectProvider } = require('./providers'); -const { CredentialTokenService } = require('./credentials'); +const { CredentialTokenService, reportRefreshFailure } = require('./credentials'); // CUSTOMER_LOAD_AND_VALIDATE_CONFIG is your startup-only configuration loader. // It must validate the complete list before any context is used. From 72fcd28453bad8c086feaf39c4d3e8b7cf94d215 Mon Sep 17 00:00:00 2001 From: Hou Chi Chan Date: Thu, 8 Oct 2026 11:23:49 -0700 Subject: [PATCH 06/14] Clarify handler edits and registered-handler feasibility evidence Preserve the existing channel-mismatch branch, specify the exact selection block to replace, separate build and transport snippets, and record the offline registered-handler check with its limitations. Co-authored-by: Copilot App <223556219+Copilot@users.noreply.github.com> Copilot-Session: 9314a109-37df-4021-8e75-bea995d57b23 --- docs/MULTI-PROVIDER-IMPLEMENTATION.md | 35 ++++++++++++++++----------- 1 file changed, 21 insertions(+), 14 deletions(-) diff --git a/docs/MULTI-PROVIDER-IMPLEMENTATION.md b/docs/MULTI-PROVIDER-IMPLEMENTATION.md index c1c05b9..482f744 100644 --- a/docs/MULTI-PROVIDER-IMPLEMENTATION.md +++ b/docs/MULTI-PROVIDER-IMPLEMENTATION.md @@ -155,15 +155,15 @@ Authenticate -> validate envelope -> decrypt -> check delivery context -> otherwise select one provider context ``` -Immediately after the existing evaluation early return, replace the single-provider lookup with -your route lookup. This fragment uses the handler's existing `fail`, `respond`, `payload`, -`correlationId` and `requestId` variables: +Immediately after the existing evaluation early return, replace the block beginning +`const provider = selectProvider(config.providerName)` through the end of `if (!provider)` with +the following. Keep the `logContext = providerContext(...)` and `provider_selected` lines that +follow it. This fragment uses the handler's existing response helpers and request variables: ```javascript const contextId = routeByChannel[payload.channelName]; const selectedContext = contexts.get(contextId); -if (!selectedContext - || selectedContext.config.providerChannel !== payload.channelName) { +if (!selectedContext) { fail('provider_selection', 'unknown_provider', 400); return respond(400, { error: 'provider_delivery_failed', correlationId, requestId, @@ -172,6 +172,7 @@ if (!selectedContext const { provider, config: providerConfig, credentials } = selectedContext; ``` +The existing channel check below this block will reject a mismatched channel in step 5. For customer/account-based routing, this is where your **verified, authorized** customer binding must select the context instead of `routeByChannel`. Evaluation must not depend on that selection or acquire provider credentials. An invalid encrypted evaluation still fails the existing checks; @@ -201,8 +202,8 @@ the `credential_unavailable` error path. For OAuth logging, pass `providerConfig `credentialContext` too. Keep the handler's [OtpDelivery](../javascript/src/functions/delivery.js) construction so phone, -message, locale and correlation handling stay unchanged. In the existing request-build and -transport `try` blocks, use the selected endpoint, adapter settings and timeout: +message, locale and correlation handling stay unchanged. In the existing **request-build** +`try` block, replace only the `provider.createRequest` call: ```javascript providerRequest = provider.createRequest({ @@ -212,8 +213,13 @@ providerRequest = provider.createRequest({ credential, env: providerConfig.env, }); +``` + +In the separate **transport** `try` block, replace only the `sendProviderRequest` call. +Keep both blocks' existing catches so build failures and transport failures retain their +different error classifications: -// Keep this in the existing transport try/catch, not the request-build block. +```javascript transportResponse = await sendProviderRequest( providerRequest, parseProviderTimeout(providerConfig.providerTimeoutMs), @@ -357,12 +363,13 @@ owns one selected cache. Update warmup/shutdown and preserve ## Scope of this guidance -An offline JavaScript prototype checked isolated Telesign API-key and Soprano OAuth contexts, -concurrent SMS dispatch, failure isolation and evaluation ordering using fake SDK/transport -boundaries. It was not a deployed or authenticated end-to-end test. Multi-provider voice, -real OAuth/Key Vault, trusted customer routing, production refresh/scale-out and live delivery -still need your implementation and validation. .NET/Python pointers are based on code inspection, -not equivalent multi-context tests. +An offline, in-memory adaptation of the registered JavaScript `SendOtp` handler checked this +SMS/Telesign and voice/Soprano design, same-provider account isolation, failures and evaluation +ordering. Its registered startup/shutdown callbacks and simulated refresh were also exercised. +Functions host registration, Azure SDK calls and provider HTTP were faked: this was not a deployed +or authenticated end-to-end test. Real OAuth/Key Vault, trusted customer routing, production +refresh/scale-out and handset delivery still need your validation. .NET/Python pointers are +based on code inspection, not equivalent multi-context tests. Runtime adapters also include Infobip and Sinch; the [setup catalog](../setup/providers/catalog.json) contains only Telesign and Soprano. Adapter availability does not imply setup coverage or account From 71d2f5d91d4a7d88ae91b09bf67a7baf93d9cfec Mon Sep 17 00:00:00 2001 From: Hou Chi Chan Date: Thu, 8 Oct 2026 15:20:54 -0700 Subject: [PATCH 07/14] Document lazy lifecycle exports and baseline test workflow Keep startup function, hook and exports consistent in the lazy variant; clarify dependency restore and adapting original singleton test fixtures. Simplify the README guide label. Co-authored-by: Copilot App <223556219+Copilot@users.noreply.github.com> Copilot-Session: 9314a109-37df-4021-8e75-bea995d57b23 --- README.md | 2 +- docs/MULTI-PROVIDER-IMPLEMENTATION.md | 37 +++++++++++++++++++++++---- 2 files changed, 33 insertions(+), 6 deletions(-) diff --git a/README.md b/README.md index e57f4f7..8fa77a7 100644 --- a/README.md +++ b/README.md @@ -14,7 +14,7 @@ For implementation details, configuration, packaging, and security behavior, see You do not need to build locally, create `local.settings.json`, or read all three language guides. For a customer-built multi-provider customization, see the -[implementation guide and offline feasibility findings](docs/MULTI-PROVIDER-IMPLEMENTATION.md). +[step-by-step implementation guide](docs/MULTI-PROVIDER-IMPLEMENTATION.md). ## Deployment options diff --git a/docs/MULTI-PROVIDER-IMPLEMENTATION.md b/docs/MULTI-PROVIDER-IMPLEMENTATION.md index 482f744..4171303 100644 --- a/docs/MULTI-PROVIDER-IMPLEMENTATION.md +++ b/docs/MULTI-PROVIDER-IMPLEMENTATION.md @@ -279,10 +279,24 @@ function stopProviderCredentialRefresh() { Keep the existing `app.hook.appStart(startProviderCredentialRefresh)` and `app.hook.appTerminate(stopProviderCredentialRefresh)` registrations once each; do not add duplicate -hooks. Build and validate `contexts` before startup runs. If you choose lazy acquisition instead, -remove the old app-start registration and its singleton warmup function, but retain the -context-closing termination hook above. In either case, no lifecycle code should continue to -resolve credentials from the original process-wide provider selection. +hooks. Build and validate `contexts` before startup runs. With prewarming, retain both functions +in the existing bottom export: + +```javascript +module.exports = { startProviderCredentialRefresh, stopProviderCredentialRefresh }; +``` + +For **lazy acquisition**, remove the app-start registration and the +`startProviderCredentialRefresh` function together. Retain the context-closing termination +function and its hook, and replace the bottom export with: + +```javascript +module.exports = { stopProviderCredentialRefresh }; +``` + +Leaving the removed startup function in `module.exports` causes a `ReferenceError` when the module +loads, before any request can run. In either variant, no lifecycle code should continue to resolve +credentials from the original process-wide provider selection. Keep cache expiry, refresh coalescing and acquisition bounds. Use a controlled worker restart when changing configuration, or implement safe draining and replacement of whole contexts. Do not @@ -302,7 +316,16 @@ Start from the existing They show how to fake credential acquisition and provider transport. Use synthetic delivery data and a locally generated encryption key; make unexpected network calls fail. -Run those existing tests from the repository root in PowerShell: +**Before editing the handler**, run the existing tests as a baseline. Use Node.js 22 and restore +missing JavaScript dependencies from the existing lockfile first. From the repository root in +PowerShell: + +```powershell +# Only if dependencies are not installed. +npm ci --prefix .\javascript --ignore-scripts --no-audit --no-fund +``` + +Then run: ```powershell node --test .\javascript\test\provider-flow.test.js ` @@ -310,6 +333,10 @@ node --test .\javascript\test\provider-flow.test.js ` .\javascript\test\sendotp.test.js ``` +**After customization**, update the test fixtures and singleton/lifecycle mocks to use your +context services before rerunning and extending these tests. The original single-provider fixtures +are not automatically valid for your customized handler. + They are a starting point, not coverage for your new router. Add tests for concurrent requests to different providers **and two accounts of the same provider**. Check the exact endpoint, credential, message and correlation ID for each request. Test unknown/duplicate/ambiguous routes, From 364d49756c9b9bc7f4e8913bbf998317ce8561df Mon Sep 17 00:00:00 2001 From: Hou Chi Chan Date: Fri, 9 Oct 2026 09:44:14 -0700 Subject: [PATCH 08/14] Clarify fixed-URL multi-provider testing and deployment boundaries Preserve deployment/restart guidance and the two-SMS testing draft. Specify shared Function App routing, Infobip base URL and sender requirements, bounded paired attempts, durable guards and no-failover/live-delivery limitations. Co-authored-by: Copilot App <223556219+Copilot@users.noreply.github.com> Copilot-Session: 9314a109-37df-4021-8e75-bea995d57b23 --- docs/MULTI-PROVIDER-IMPLEMENTATION.md | 134 +++++++++++++++++++++++++- 1 file changed, 132 insertions(+), 2 deletions(-) diff --git a/docs/MULTI-PROVIDER-IMPLEMENTATION.md b/docs/MULTI-PROVIDER-IMPLEMENTATION.md index 4171303..89e5f03 100644 --- a/docs/MULTI-PROVIDER-IMPLEMENTATION.md +++ b/docs/MULTI-PROVIDER-IMPLEMENTATION.md @@ -15,7 +15,9 @@ Anything named `CUSTOMER_*` is a placeholder you must implement; it is not a rep ## 1. Choose a routing rule you control -Start with a simple, server-owned rule. For example, use one account for SMS and another for voice: +Start with a simple, server-owned rule. The steps below use an **illustrative channel policy**: +one account for SMS and another for voice. This is different from the +[two-SMS fixed-URL test](#testing-two-sms-providers-in-one-deployment) later in this guide. ```javascript // Module scope in your customized SendOtp.js, not supplied by the caller. @@ -362,6 +364,131 @@ and attempt limits, and a **one-attempt safety gate**. Provider acceptance is no delivery. Plan an explicit rollout and rollback; do not hide failures behind automatic fallback. Nothing in this walkthrough authorizes a SAS trigger or provider send. +### What needs deployment, configuration, or restart? + +Deploy your customized application code once, after its offline tests pass. Adding provider +settings to an unchanged checkout does not implement this guide: the shipped handler still +selects one provider per deployment. A customer implementation can host multiple account +contexts in **one Function App and one code package**, with multiple HTTP-triggered functions +sharing handler code. This is not one deployment per provider. Separate deployments are needed +only when your security or operational isolation requirements call for them. + +| Change | Required action | +| --- | --- | +| Add or change the router, an adapter, a Function entry point, or a code-owned endpoint allowlist | Build, test and deploy a new application package. | +| Change account references, scopes, identities, or routing in an already implemented startup configuration loader | Validate the complete configuration, apply it, and perform a controlled worker restart. If configuration is embedded in code, deploy a new package instead. | +| Rotate a referenced secret | Update the approved secret and verify credential-cache refresh. Updating the vault alone is not proof that running workers use the new version; use a controlled restart when needed. Update pinned secret-version references explicitly. | +| Repeat an authorized test with unchanged code and configuration | No redeployment is needed. Verify the deployed package identifier, routing, authentication and credential readiness first. | +| Update only this documentation | No runtime deployment is needed. | + +The existing deployment automation does not create your custom account registry, routes or +configuration loader. Adapt your deployment process to provision only the required identities, +vault permissions and provider consent, then deploy the same reviewed package and validated +configuration to each intended instance. Preserve inbound authentication and decryption settings. +Treat code and configuration as one rollout, and retain their previous versions for rollback. + +### Testing two SMS providers in one deployment + +The channel example earlier distinguishes SMS/Telesign from voice/Soprano. It cannot choose +between two SMS providers. For the reported customer test design, use these fixed bindings in +**one Function App and one package**: + +```text +POST /api/SendOtp -> telesign-primary -> Telesign -> channel=sms +POST /api/SendOtpInfobip -> infobip-primary -> Infobip -> channel=sms +``` + +The caller chooses the URL; server code binds that URL to its approved provider context. +No provider-selection header is needed. Authorize the caller for the account behind each entry +point: the URL's existence alone is not account authorization. + +In your customization of [SendOtp.js](../javascript/src/functions/SendOtp.js), extract the common +callback into a handler factory that captures a fixed context ID. Register two HTTP-triggered +functions using the existing `app.http` pattern, with explicit `route: 'SendOtp'` and +`route: 'SendOtpInfobip'` (assuming the standard `/api` route prefix). Register each name/route +once; replace the original registration rather than adding a duplicate. The shared handler keeps +the same validation, decryption, evaluation early return and error/response mapping. Only after +evaluation returns does it look up its captured ID, instead of using `routeByChannel`. Both +contexts must require `sms`. A handler factory is **code you implement**, not a shipped +`createHandler` API. Neither payload fields nor headers may override the captured binding. + +Keep the Telesign configuration from step 2 and add this entry through your custom loader: + +```json +{ + "id": "infobip-primary", + "settings": { + "EPP_PROVIDER_NAME": "infobip", + "EPP_PROVIDER_CHANNEL": "sms", + "EPP_PROVIDER_AUTH_MODE": "apiKey", + "EPP_PROVIDER_ENDPOINT": "https://", + "EPP_PROVIDER_ACCOUNT_NAME": "Verify", + "EPP_PROVIDER_TIMEOUT_MS": "1500", + "KEY_VAULT_URL": "https://.vault.azure.net", + "AZURE_CLIENT_ID": "" + } +} +``` + +The [Infobip adapter](../javascript/src/functions/providers/infobip.js) reads `infobip-api-key` +and appends `/sms/3/messages` to the configured **base URL**. Do not put that operation suffix +in `EPP_PROVIDER_ENDPOINT`: it must be appended exactly once. In your startup loader, validate +the approved HTTPS base host with no operation path, query or fragment, and either reject a +trailing slash or remove it before freezing the settings; otherwise the adapter creates a double +separator. `Verify` was the sender setting for this test design, not a universally valid sender. +Use the sender approved for your Infobip account and destination; do not rely on the adapter's +`Verify` default as evidence of approval. + +Create a separate long-lived credential service/cache for each account, with the required vault +permissions and managed identity access. Complete each provider's account, route, sender and +recipient prerequisites. Keep shared inbound authentication/decryption separate from these +outbound credentials. The existing setup catalog does not provision this Infobip customization. + +This is **provider selection, not automatic provider failover**. The test caller chooses a fixed +entry point; its server-owned binding chooses the account. Both requests still use the validated +`sms` channel, and no provider-selection header is needed. A failed or timed-out request must not +silently switch accounts, retry or send through both. **SAS calling one configured URL keeps +using that URL**; adding another function does not make SAS select it. Single-URL provider +selection or automatic fallback requires a separately designed trusted policy and handling for +delivery uncertainty and duplicates. [Front Door regional failover](FRONTDOOR.md) is separate +from provider selection. + +Only under a separately approved test plan, after authorized encrypted evaluation and rejection +checks pass on both entry points, a bounded comparison can use **five paired rounds**: +one Telesign message and one Infobip message per round, **ten planned message attempts if all +five pairs finish**. Agree on that limit and an approved recipient before running. These labels +describe planned attempts, not verified sends or delivery: + +```text +EPP multi-provider test: Telesign - attempt 1/5. +EPP multi-provider test: Infobip - attempt 1/5. +... +EPP multi-provider test: Telesign - attempt 5/5. +EPP multi-provider test: Infobip - attempt 5/5. +``` + +Alternate providers while keeping the deployed package and account configuration unchanged. +Assign a unique correlation ID to every request. Atomically persist a durable guard recording +intent before dispatch, keyed by test run, provider and attempt number; refuse an already-recorded +attempt, including after a runner restart. This guard is **not provider idempotency or proof of +exactly-once delivery**: a crash or timeout can leave the outcome uncertain. +Do not retry an uncertain send automatically. Stop on unexpected failures, report actual attempted +counts and partial results, and obtain separate approval for any replacement attempt. Do not +automatically top up an aborted run to ten or turn errors into fallback sends. + +For each planned request, record its provider, round, package identifier, endpoint response, +nonce validation, latency and correlated provider-dispatch evidence. Verify exactly one dispatch +to the intended provider and no cross-account fallback. Record provider acceptance and handset +receipt separately; `PENDING` or HTTP 200 is not a delivery receipt. Keep phone numbers, message +contents, tokens and credentials out of telemetry. + +A sequential comparison does not prove concurrency, capacity, automatic provider failover, +Front Door regional resilience or real SAS integration. Preserve the existing inbound authentication +and origin/network restrictions for both functions; use only the approved test access path. If that +path is unavailable, stop and arrange authorized access rather than opening ingress or bypassing +restrictions. This section documents a custom test design, not authorization to execute it or a +claim that all planned messages were accepted or delivered. + ## Using .NET or Python instead **.NET:** Start with [SendOtp.cs](../dotnet/Functions/SendOtp.cs) and @@ -400,4 +527,7 @@ based on code inspection, not equivalent multi-context tests. Runtime adapters also include Infobip and Sinch; the [setup catalog](../setup/providers/catalog.json) contains only Telesign and Soprano. Adapter availability does not imply setup coverage or account -entitlement. No Azure, Graph, Key Vault, SAS or provider operations were performed for this guide. +entitlement. The original offline feasibility check performed no Azure, Graph, Key Vault, SAS or +provider operations. Report any subsequent live test of a customer customization separately, +including its deployed revision, authorized workload, provider responses and delivery limitations. +Such a test does not make the custom router part of the shipped sample. From 4ac6ca0123f5855a0664c1e0688c0b4707bca4e9 Mon Sep 17 00:00:00 2001 From: Hou Chi Chan Date: Fri, 9 Oct 2026 09:45:56 -0700 Subject: [PATCH 09/14] Separate fixed-route test evidence and upstream routing requirements Clarify that earlier channel-based concurrency checks do not cover the Telesign/Infobip fixed-route deployment, and that upstream path forwarding must be configured without weakening authentication. Co-authored-by: Copilot App <223556219+Copilot@users.noreply.github.com> Copilot-Session: 9314a109-37df-4021-8e75-bea995d57b23 --- docs/MULTI-PROVIDER-IMPLEMENTATION.md | 9 +++++++++ 1 file changed, 9 insertions(+) diff --git a/docs/MULTI-PROVIDER-IMPLEMENTATION.md b/docs/MULTI-PROVIDER-IMPLEMENTATION.md index 89e5f03..8acc7e0 100644 --- a/docs/MULTI-PROVIDER-IMPLEMENTATION.md +++ b/docs/MULTI-PROVIDER-IMPLEMENTATION.md @@ -402,6 +402,10 @@ The caller chooses the URL; server code binds that URL to its approved provider No provider-selection header is needed. Authorize the caller for the account behind each entry point: the URL's existence alone is not account authorization. +If these URLs are behind a proxy or Front Door, configure and verify forwarding for both intended +paths while retaining authentication and origin restrictions. Adding a Function route does not +automatically update upstream routing configuration. + In your customization of [SendOtp.js](../javascript/src/functions/SendOtp.js), extract the common callback into a handler factory that captures a fixed context ID. Register two HTTP-triggered functions using the existing `app.http` pattern, with explicit `route: 'SendOtp'` and @@ -489,6 +493,11 @@ path is unavailable, stop and arrange authorized access rather than opening ingr restrictions. This section documents a custom test design, not authorization to execute it or a claim that all planned messages were accepted or delivered. +Keep the evidence separate: the earlier offline registered-handler checks exercised channel-based +Telesign SMS/Soprano voice routing, not this Telesign/Infobip pair of fixed HTTP routes. Their +concurrency results do not validate the new deployment or extend the user's reported sequential +URL-based test into a concurrency test. + ## Using .NET or Python instead **.NET:** Start with [SendOtp.cs](../dotnet/Functions/SendOtp.cs) and From 6eef42b9a18e6f68f7d986af9b6de442f7a0e393 Mon Sep 17 00:00:00 2001 From: Hou Chi Chan Date: Fri, 9 Oct 2026 09:48:24 -0700 Subject: [PATCH 10/14] Clarify identity attachment and scoped vault access prerequisites Explain user-assigned and system-assigned identity permissions alongside per-account credential configuration, with an onboarding reference. Co-authored-by: Copilot App <223556219+Copilot@users.noreply.github.com> Copilot-Session: 9314a109-37df-4021-8e75-bea995d57b23 --- docs/MULTI-PROVIDER-IMPLEMENTATION.md | 8 +++++++- 1 file changed, 7 insertions(+), 1 deletion(-) diff --git a/docs/MULTI-PROVIDER-IMPLEMENTATION.md b/docs/MULTI-PROVIDER-IMPLEMENTATION.md index 8acc7e0..172b6fe 100644 --- a/docs/MULTI-PROVIDER-IMPLEMENTATION.md +++ b/docs/MULTI-PROVIDER-IMPLEMENTATION.md @@ -444,7 +444,13 @@ Use the sender approved for your Infobip account and destination; do not rely on `Verify` default as evidence of approval. Create a separate long-lived credential service/cache for each account, with the required vault -permissions and managed identity access. Complete each provider's account, route, sender and +permissions and managed identity access. Attach each referenced user-assigned managed identity +to the Function App and grant the vault-reading identity scoped secret-read access to its intended +vault, such as **Key Vault Secrets User** under RBAC or the approved access-policy equivalent. +If using the Function App's system-assigned identity, grant that identity instead. Context IDs and +managed-identity client-ID settings do not attach identities or grant permissions. See the +[provider credential onboarding guidance](ONBOARDING.md#complete-provider-authentication). +Complete each provider's account, route, sender and recipient prerequisites. Keep shared inbound authentication/decryption separate from these outbound credentials. The existing setup catalog does not provision this Infobip customization. From 8f3ee9f0b23b0c7bd41850ccd9b0c0a57d1d985e Mon Sep 17 00:00:00 2001 From: Hou Chi Chan Date: Fri, 9 Oct 2026 09:54:04 -0700 Subject: [PATCH 11/14] Record test-owner confirmation of SMS receipt Attribute receipt to the reported sequential test owner's confirmation, separate from provider statuses and unverified per-attempt evidence. Preserve all rollout and test limitations. Co-authored-by: Copilot App <223556219+Copilot@users.noreply.github.com> Copilot-Session: 9314a109-37df-4021-8e75-bea995d57b23 --- docs/MULTI-PROVIDER-IMPLEMENTATION.md | 10 ++++++++-- 1 file changed, 8 insertions(+), 2 deletions(-) diff --git a/docs/MULTI-PROVIDER-IMPLEMENTATION.md b/docs/MULTI-PROVIDER-IMPLEMENTATION.md index 172b6fe..dbb4ef5 100644 --- a/docs/MULTI-PROVIDER-IMPLEMENTATION.md +++ b/docs/MULTI-PROVIDER-IMPLEMENTATION.md @@ -492,12 +492,18 @@ to the intended provider and no cross-account fallback. Record provider acceptan receipt separately; `PENDING` or HTTP 200 is not a delivery receipt. Keep phone numbers, message contents, tokens and credentials out of telemetry. +**Test-owner-confirmed handset receipt:** the test owner reports that all SMS in the reported +sequential fixed-URL Telesign/Infobip test were received. This confirmation comes from the owner, +not from HTTP 200, provider acceptance or `PENDING` statuses. It does not establish exact +per-provider/per-attempt counts, timestamps, provider delivery receipts or independently verified +provider attribution. + A sequential comparison does not prove concurrency, capacity, automatic provider failover, Front Door regional resilience or real SAS integration. Preserve the existing inbound authentication and origin/network restrictions for both functions; use only the approved test access path. If that path is unavailable, stop and arrange authorized access rather than opening ingress or bypassing -restrictions. This section documents a custom test design, not authorization to execute it or a -claim that all planned messages were accepted or delivered. +restrictions. This section documents a custom test design and the owner's reported receipt, +not authorization for another test or an independently verified accounting of every planned attempt. Keep the evidence separate: the earlier offline registered-handler checks exercised channel-based Telesign SMS/Soprano voice routing, not this Telesign/Infobip pair of fixed HTTP routes. Their From c76dc000ae432ff50a548e92c6bc3991e69b38dc Mon Sep 17 00:00:00 2001 From: Hou Chi Chan Date: Fri, 9 Oct 2026 10:21:09 -0700 Subject: [PATCH 12/14] Constrain multi-provider guide to one SAS-facing URL Use a fixed startup-owned provider binding per validated channel. Remove dual-endpoint routing and paired test history, and describe Infobip only as an administratively deployed SMS replacement with isolated credentials and no fallback. Co-authored-by: Copilot App <223556219+Copilot@users.noreply.github.com> Copilot-Session: 9314a109-37df-4021-8e75-bea995d57b23 --- docs/MULTI-PROVIDER-IMPLEMENTATION.md | 191 +++++++++----------------- 1 file changed, 67 insertions(+), 124 deletions(-) diff --git a/docs/MULTI-PROVIDER-IMPLEMENTATION.md b/docs/MULTI-PROVIDER-IMPLEMENTATION.md index dbb4ef5..3bb9944 100644 --- a/docs/MULTI-PROVIDER-IMPLEMENTATION.md +++ b/docs/MULTI-PROVIDER-IMPLEMENTATION.md @@ -1,23 +1,22 @@ # Add multiple provider configurations to your application -You can adapt this sample to choose between multiple provider accounts in one application. -The key is to give each account its own configuration and credential cache, then select exactly -one account for each request. +You can adapt this sample to use multiple provider accounts behind **one SAS-facing URL: +`/api/SendOtp`**. SAS keeps calling its configured URL. The server maps the validated channel to +exactly one active provider/account context: Telesign for SMS and Soprano for voice in this +walkthrough. Each context has its own configuration and credential cache. The sample does **not** do this out of the box: it still uses one provider per deployment through `EPP_PROVIDER_NAME`. This walkthrough shows where to make the changes in your own JavaScript -application. It does not add a router or setup script to the repository. If accounts need a strong -security boundary, keep them in separate deployments instead. +application. It does not add a router or setup script to the repository. The snippets are **integration examples, not a drop-in handler**. They fit into [SendOtp.js](../javascript/src/functions/SendOtp.js) and use its existing variables and helpers. Anything named `CUSTOMER_*` is a placeholder you must implement; it is not a repository API. -## 1. Choose a routing rule you control +## 1. Bind each channel to one provider at startup -Start with a simple, server-owned rule. The steps below use an **illustrative channel policy**: -one account for SMS and another for voice. This is different from the -[two-SMS fixed-URL test](#testing-two-sms-providers-in-one-deployment) later in this guide. +Keep the existing `SendOtp` HTTP-trigger registration and `/api/SendOtp` path. Define this +server-owned map once at startup; do not add another SAS endpoint or a caller-selected provider: ```javascript // Module scope in your customized SendOtp.js, not supplied by the caller. @@ -31,11 +30,12 @@ Open [entraPayload.js](../javascript/src/functions/entraPayload.js) to see how t validates the channel and exposes `payload.channelName`. Use that validated value with your approved policy. Reject unsupported or ambiguous routes; do not silently choose a default. -If you need two accounts for the same channel, first decide how your application will verify -the customer or endpoint identity that owns the request. The existing contract has no trusted -provider-selector field. An arbitrary header, `providerId`, correlation ID or unverified body -`tenantId` is **not** authority to select an account. The SAS caller identity may not distinguish -your customers either. Keep endpoint URLs, vaults, scopes and secret names under server control. +Each supported channel has exactly one active account. Two SMS providers cannot both be +caller-selected in this design. Changing the active SMS provider is an administrative configuration +change, validated and applied through deployment or a controlled restart as described below. +Do not use headers, `providerId`, correlation IDs or body `tenantId` to override the map. There is +no random or round-robin selection, automatic switching, fallback, fan-out or resend. Keep provider +endpoint URLs, vaults, scopes and secret names under server control. ## 2. Define one configuration per provider account @@ -175,10 +175,9 @@ const { provider, config: providerConfig, credentials } = selectedContext; ``` The existing channel check below this block will reject a mismatched channel in step 5. -For customer/account-based routing, this is where your **verified, authorized** customer binding -must select the context instead of `routeByChannel`. Evaluation must not depend on that selection -or acquire provider credentials. An invalid encrypted evaluation still fails the existing checks; -do not add a shortcut that returns a nonce before decryption. +Use only the startup-owned `routeByChannel` map for selection. Evaluation must not depend on +that selection or acquire provider credentials. An invalid encrypted evaluation still fails the +existing checks; do not add a shortcut that returns a nonce before decryption. ## 5. Use the selected context for one dispatch @@ -301,13 +300,12 @@ loads, before any request can run. In either variant, no lifecycle code should c credentials from the original process-wide provider selection. Keep cache expiry, refresh coalescing and acquisition bounds. Use a controlled worker restart when -changing configuration, or implement safe draining and replacement of whole contexts. Do not +changing startup-loaded configuration, rebuilding each context and its cache. Do not modify a live service's inputs. Evaluation requests must not start credential acquisition, but already-running background refresh is independent of a request's evaluation branch. Separate contexts provide logical isolation, not a tenant security boundary. A compromised worker -may reach every identity assigned to it. Review least privilege and use separate deployments when -that risk is unacceptable. +may reach every identity assigned to it. Review least privilege before approving this design. ## 7. Test your changes offline first @@ -340,7 +338,8 @@ context services before rerunning and extending these tests. The original single are not automatically valid for your customized handler. They are a starting point, not coverage for your new router. Add tests for concurrent requests to -different providers **and two accounts of the same provider**. Check the exact endpoint, +different channel bindings, including separate SMS/voice test accounts of the same adapter. +Do not introduce two active SMS choices to run these tests. Check the exact endpoint, credential, message and correlation ID for each request. Test unknown/duplicate/ambiguous routes, channel mismatch, expired or failing credentials, provider errors, timeouts, malformed responses, shutdown and configuration replacement. Failures must cause no second dispatch. @@ -368,55 +367,32 @@ Nothing in this walkthrough authorizes a SAS trigger or provider send. Deploy your customized application code once, after its offline tests pass. Adding provider settings to an unchanged checkout does not implement this guide: the shipped handler still -selects one provider per deployment. A customer implementation can host multiple account -contexts in **one Function App and one code package**, with multiple HTTP-triggered functions -sharing handler code. This is not one deployment per provider. Separate deployments are needed -only when your security or operational isolation requirements call for them. +selects one provider per deployment. This customization hosts the SMS and voice account contexts +in **one Function App and one code package**, behind the existing `SendOtp` HTTP-triggered function. +It is not one deployment per provider, and it does not change the SAS-facing URL. | Change | Required action | | --- | --- | -| Add or change the router, an adapter, a Function entry point, or a code-owned endpoint allowlist | Build, test and deploy a new application package. | -| Change account references, scopes, identities, or routing in an already implemented startup configuration loader | Validate the complete configuration, apply it, and perform a controlled worker restart. If configuration is embedded in code, deploy a new package instead. | +| Change the handler, channel map, an adapter, or a code-owned endpoint allowlist | Build, test and deploy a new application package. | +| Change account references, scopes or identities in the startup configuration loader | Validate the complete configuration, apply it, and perform a controlled worker restart. If configuration is embedded in code, deploy a new package instead. | | Rotate a referenced secret | Update the approved secret and verify credential-cache refresh. Updating the vault alone is not proof that running workers use the new version; use a controlled restart when needed. Update pinned secret-version references explicitly. | | Repeat an authorized test with unchanged code and configuration | No redeployment is needed. Verify the deployed package identifier, routing, authentication and credential readiness first. | | Update only this documentation | No runtime deployment is needed. | -The existing deployment automation does not create your custom account registry, routes or +The existing deployment automation does not create your custom account contexts, channel map or configuration loader. Adapt your deployment process to provision only the required identities, vault permissions and provider consent, then deploy the same reviewed package and validated configuration to each intended instance. Preserve inbound authentication and decryption settings. Treat code and configuration as one rollout, and retain their previous versions for rollback. -### Testing two SMS providers in one deployment +### Replace the active SMS provider with Infobip -The channel example earlier distinguishes SMS/Telesign from voice/Soprano. It cannot choose -between two SMS providers. For the reported customer test design, use these fixed bindings in -**one Function App and one package**: +This is an optional administrative replacement of Telesign for SMS, not an additional caller +choice. `/api/SendOtp` and the Soprano voice binding remain unchanged. -```text -POST /api/SendOtp -> telesign-primary -> Telesign -> channel=sms -POST /api/SendOtpInfobip -> infobip-primary -> Infobip -> channel=sms -``` - -The caller chooses the URL; server code binds that URL to its approved provider context. -No provider-selection header is needed. Authorize the caller for the account behind each entry -point: the URL's existence alone is not account authorization. - -If these URLs are behind a proxy or Front Door, configure and verify forwarding for both intended -paths while retaining authentication and origin restrictions. Adding a Function route does not -automatically update upstream routing configuration. - -In your customization of [SendOtp.js](../javascript/src/functions/SendOtp.js), extract the common -callback into a handler factory that captures a fixed context ID. Register two HTTP-triggered -functions using the existing `app.http` pattern, with explicit `route: 'SendOtp'` and -`route: 'SendOtpInfobip'` (assuming the standard `/api` route prefix). Register each name/route -once; replace the original registration rather than adding a duplicate. The shared handler keeps -the same validation, decryption, evaluation early return and error/response mapping. Only after -evaluation returns does it look up its captured ID, instead of using `routeByChannel`. Both -contexts must require `sms`. A handler factory is **code you implement**, not a shipped -`createHandler` API. Neither payload fields nor headers may override the captured binding. - -Keep the Telesign configuration from step 2 and add this entry through your custom loader: +1. Replace the `telesign-primary` entry in your custom loader's active configuration with the + following `infobip-primary` entry. Keep the Soprano entry. Supply approved account values, + never secret values: ```json { @@ -426,7 +402,7 @@ Keep the Telesign configuration from step 2 and add this entry through your cust "EPP_PROVIDER_CHANNEL": "sms", "EPP_PROVIDER_AUTH_MODE": "apiKey", "EPP_PROVIDER_ENDPOINT": "https://", - "EPP_PROVIDER_ACCOUNT_NAME": "Verify", + "EPP_PROVIDER_ACCOUNT_NAME": "", "EPP_PROVIDER_TIMEOUT_MS": "1500", "KEY_VAULT_URL": "https://.vault.azure.net", "AZURE_CLIENT_ID": "" @@ -439,76 +415,43 @@ and appends `/sms/3/messages` to the configured **base URL**. Do not put that op in `EPP_PROVIDER_ENDPOINT`: it must be appended exactly once. In your startup loader, validate the approved HTTPS base host with no operation path, query or fragment, and either reject a trailing slash or remove it before freezing the settings; otherwise the adapter creates a double -separator. `Verify` was the sender setting for this test design, not a universally valid sender. -Use the sender approved for your Infobip account and destination; do not rely on the adapter's -`Verify` default as evidence of approval. +separator. The adapter defaults to `Verify` if the sender setting is absent; that is not a +universally valid sender. Set the sender approved for your Infobip account and destination. -Create a separate long-lived credential service/cache for each account, with the required vault -permissions and managed identity access. Attach each referenced user-assigned managed identity -to the Function App and grant the vault-reading identity scoped secret-read access to its intended +2. Complete the account, sender, destination and vault-access prerequisites before activating + the replacement. Create a new long-lived credential service/cache for the Infobip context; + never reuse the Telesign cache with different inputs. + +Attach each referenced user-assigned managed identity to the Function App and grant the +vault-reading identity scoped secret-read access to its intended vault, such as **Key Vault Secrets User** under RBAC or the approved access-policy equivalent. If using the Function App's system-assigned identity, grant that identity instead. Context IDs and managed-identity client-ID settings do not attach identities or grant permissions. See the [provider credential onboarding guidance](ONBOARDING.md#complete-provider-authentication). -Complete each provider's account, route, sender and -recipient prerequisites. Keep shared inbound authentication/decryption separate from these -outbound credentials. The existing setup catalog does not provision this Infobip customization. - -This is **provider selection, not automatic provider failover**. The test caller chooses a fixed -entry point; its server-owned binding chooses the account. Both requests still use the validated -`sms` channel, and no provider-selection header is needed. A failed or timed-out request must not -silently switch accounts, retry or send through both. **SAS calling one configured URL keeps -using that URL**; adding another function does not make SAS select it. Single-URL provider -selection or automatic fallback requires a separately designed trusted policy and handling for -delivery uncertainty and duplicates. [Front Door regional failover](FRONTDOOR.md) is separate -from provider selection. - -Only under a separately approved test plan, after authorized encrypted evaluation and rejection -checks pass on both entry points, a bounded comparison can use **five paired rounds**: -one Telesign message and one Infobip message per round, **ten planned message attempts if all -five pairs finish**. Agree on that limit and an approved recipient before running. These labels -describe planned attempts, not verified sends or delivery: - -```text -EPP multi-provider test: Telesign - attempt 1/5. -EPP multi-provider test: Infobip - attempt 1/5. -... -EPP multi-provider test: Telesign - attempt 5/5. -EPP multi-provider test: Infobip - attempt 5/5. -``` - -Alternate providers while keeping the deployed package and account configuration unchanged. -Assign a unique correlation ID to every request. Atomically persist a durable guard recording -intent before dispatch, keyed by test run, provider and attempt number; refuse an already-recorded -attempt, including after a runner restart. This guard is **not provider idempotency or proof of -exactly-once delivery**: a crash or timeout can leave the outcome uncertain. -Do not retry an uncertain send automatically. Stop on unexpected failures, report actual attempted -counts and partial results, and obtain separate approval for any replacement attempt. Do not -automatically top up an aborted run to ten or turn errors into fallback sends. - -For each planned request, record its provider, round, package identifier, endpoint response, -nonce validation, latency and correlated provider-dispatch evidence. Verify exactly one dispatch -to the intended provider and no cross-account fallback. Record provider acceptance and handset -receipt separately; `PENDING` or HTTP 200 is not a delivery receipt. Keep phone numbers, message -contents, tokens and credentials out of telemetry. - -**Test-owner-confirmed handset receipt:** the test owner reports that all SMS in the reported -sequential fixed-URL Telesign/Infobip test were received. This confirmation comes from the owner, -not from HTTP 200, provider acceptance or `PENDING` statuses. It does not establish exact -per-provider/per-attempt counts, timestamps, provider delivery receipts or independently verified -provider attribution. - -A sequential comparison does not prove concurrency, capacity, automatic provider failover, -Front Door regional resilience or real SAS integration. Preserve the existing inbound authentication -and origin/network restrictions for both functions; use only the approved test access path. If that -path is unavailable, stop and arrange authorized access rather than opening ingress or bypassing -restrictions. This section documents a custom test design and the owner's reported receipt, -not authorization for another test or an independently verified accounting of every planned attempt. - -Keep the evidence separate: the earlier offline registered-handler checks exercised channel-based -Telesign SMS/Soprano voice routing, not this Telesign/Infobip pair of fixed HTTP routes. Their -concurrency results do not validate the new deployment or extend the user's reported sequential -URL-based test into a concurrency test. +Keep shared inbound authentication/decryption separate from these outbound credentials. +The existing setup catalog does not provision this Infobip customization. + +3. Change only the `sms` value in `routeByChannel` from `telesign-primary` to `infobip-primary`. + Validate that each supported channel resolves to one configured context with the matching + channel/authentication mode and an approved provider endpoint. Reject missing or duplicate + configuration, and rerun the offline handler tests with the replacement binding. +4. The map in section 1 is code-owned, so build and deploy the reviewed package with the new map + and active configuration. For later changes to values already read by the startup loader, + validate and apply the complete configuration with a controlled worker restart. Close old + contexts and initialize fresh caches; do not switch an in-flight request or retry its delivery. + Verify the deployed package/configuration version before any separately authorized live check. + +SAS still sends to `/api/SendOtp`; it does not select the replacement provider. An SMS failure +must not return to Telesign, retry or send through both providers. [Front Door regional +failover](FRONTDOOR.md) remains separate: retain forwarding for `/api/SendOtp` and preserve +inbound authentication and origin/network restrictions. + +For an authorized live check, record unique correlations and durable intent-before-send attempt +guards. Stop on unexpected failures and do not retry uncertain sends automatically. A guard is +not provider idempotency or proof of exactly-once delivery. Record actual attempts, provider +acceptance and handset receipt separately: HTTP 200 or `PENDING` alone is not handset proof. +A sequential check does not establish concurrency, capacity, automatic failover, regional +resilience or real SAS integration. This replacement procedure authorizes no live operation. ## Using .NET or Python instead @@ -542,7 +485,7 @@ An offline, in-memory adaptation of the registered JavaScript `SendOtp` handler SMS/Telesign and voice/Soprano design, same-provider account isolation, failures and evaluation ordering. Its registered startup/shutdown callbacks and simulated refresh were also exercised. Functions host registration, Azure SDK calls and provider HTTP were faked: this was not a deployed -or authenticated end-to-end test. Real OAuth/Key Vault, trusted customer routing, production +or authenticated end-to-end test. Real OAuth/Key Vault, deployed authentication, production refresh/scale-out and handset delivery still need your validation. .NET/Python pointers are based on code inspection, not equivalent multi-context tests. From 9bbf48839e5a6045fd499221578aaabec3d4fb6a Mon Sep 17 00:00:00 2001 From: Hou Chi Chan Date: Fri, 9 Oct 2026 12:59:00 -0700 Subject: [PATCH 13/14] Refocus guide on guarded primary-to-secondary provider fallback Document same-channel primary/secondary accounts behind one SAS URL, deny-by-default nonacceptance policy, durable operation guards and aggregate deadlines. Remove the per-channel provider split and distinguish new policy-model evidence from prior routing checks. Co-authored-by: Copilot App <223556219+Copilot@users.noreply.github.com> Copilot-Session: 9314a109-37df-4021-8e75-bea995d57b23 --- docs/MULTI-PROVIDER-IMPLEMENTATION.md | 669 ++++++++++---------------- 1 file changed, 252 insertions(+), 417 deletions(-) diff --git a/docs/MULTI-PROVIDER-IMPLEMENTATION.md b/docs/MULTI-PROVIDER-IMPLEMENTATION.md index 3bb9944..4a8f4a4 100644 --- a/docs/MULTI-PROVIDER-IMPLEMENTATION.md +++ b/docs/MULTI-PROVIDER-IMPLEMENTATION.md @@ -1,52 +1,52 @@ -# Add multiple provider configurations to your application +# Add conservative primary-to-secondary provider fallback -You can adapt this sample to use multiple provider accounts behind **one SAS-facing URL: -`/api/SendOtp`**. SAS keeps calling its configured URL. The server maps the validated channel to -exactly one active provider/account context: Telesign for SMS and Soprano for voice in this -walkthrough. Each context has its own configuration and credential cache. +You can adapt this application to try a secondary provider **only when the primary is known +not to have accepted the message**. Keep one SAS-facing URL, `/api/SendOtp`, and use Telesign +as the SMS primary and Infobip as the SMS secondary. Both serve the same channel; the caller +does not select an account or supply a provider header. -The sample does **not** do this out of the box: it still uses one provider per deployment through -`EPP_PROVIDER_NAME`. This walkthrough shows where to make the changes in your own JavaScript -application. It does not add a router or setup script to the repository. +The sample still ships with one provider per deployment and **no provider fallback**. This guide +describes customer code you must implement and review. A timeout, lost response or generic +provider error is not proof of nonacceptance: in those cases, do not automatically send again. -The snippets are **integration examples, not a drop-in handler**. They fit into -[SendOtp.js](../javascript/src/functions/SendOtp.js) and use its existing variables and helpers. -Anything named `CUSTOMER_*` is a placeholder you must implement; it is not a repository API. +## 1. Set a fixed provider order and a deny-by-default policy -## 1. Bind each channel to one provider at startup - -Keep the existing `SendOtp` HTTP-trigger registration and `/api/SendOtp` path. Define this -server-owned map once at startup; do not add another SAS endpoint or a caller-selected provider: +Keep the existing HTTP registration in [SendOtp.js](../javascript/src/functions/SendOtp.js). +Define the provider order in server-owned startup configuration: ```javascript -// Module scope in your customized SendOtp.js, not supplied by the caller. -const routeByChannel = Object.freeze({ - sms: 'telesign-primary', - voice: 'soprano-primary', +const providerOrder = Object.freeze({ + channel: 'sms', + primary: 'telesign-primary', + secondary: 'infobip-secondary', }); ``` -Open [entraPayload.js](../javascript/src/functions/entraPayload.js) to see how the handler -validates the channel and exposes `payload.channelName`. Use that validated value with your -approved policy. Reject unsupported or ambiguous routes; do not silently choose a default. +Reject unsupported channels and invalid configuration. Do not accept provider IDs, endpoints, +vaults or scopes from request headers or bodies. There is no round-robin selection, fan-out, +recursive fallback or primary retry. The secondary can be attempted at most once. -Each supported channel has exactly one active account. Two SMS providers cannot both be -caller-selected in this design. Changing the active SMS provider is an administrative configuration -change, validated and applied through deployment or a controlled restart as described below. -Do not use headers, `providerId`, correlation IDs or body `tenantId` to override the map. There is -no random or round-robin selection, automatic switching, fallback, fan-out or resend. Keep provider -endpoint URLs, vaults, scopes and secret names under server control. +Keep **routing order** separate from **fallback eligibility**. Use this policy: -## 2. Define one configuration per provider account +| Primary result | Action | +| --- | --- | +| Accepted, successful, or `PENDING` | Stop. Do not contact the secondary. Acceptance is not handset delivery. | +| Block/fraud denial, bad recipient or payload, invalid inbound authentication | Stop. Never use another provider to bypass the denial or validation. | +| Unknown status, malformed/lost response, timeout, connection/read error, or cancellation after dispatch | Record a terminal or uncertain outcome. Do not automatically contact the secondary. | +| Generic 4xx/5xx, `Fail`, or `provider_rejected` | No fallback based on that classification alone. | +| Primary-only credential acquisition fails before transport is entered | Eligible only if this pre-dispatch fact is recorded, the account policy explicitly permits it, the durable operation guard allows the transition, and the request budget is sufficient. | +| Provider-specific, documented proof that the primary did not accept the operation | Eligible only through a reviewed nonacceptance classifier, explicit account policy, the durable guard and sufficient budget. | -Open [config.js](../javascript/src/functions/config.js). Its `readConfig(env)` function accepts -an explicit settings object, so you can reuse it without changing `process.env` per request. +**Default to no fallback.** No real Telesign or Infobip response is designated safe to fall back +from by this guide. Obtain and review provider-specific semantics before enabling that path. +An expired/rejected inbound token never reaches it. A provider credential failure must not bypass +account suspension, consent requirements, entitlement restrictions or a provider block. -Give each account a unique context ID. For example, two Telesign accounts would need different -IDs even though both have `EPP_PROVIDER_NAME: "telesign"`. +## 2. Configure and prevalidate both accounts -Use the following as a starting point for **your own configuration format**. Replace the -placeholders with approved values. The sample does not load this JSON automatically. +[config.js](../javascript/src/functions/config.js) accepts `readConfig(settings)`, so each +account can use its own immutable settings without changing `process.env` per request. +The following is a **customer-owned configuration format**, not a file the sample loads: ```json [ @@ -57,441 +57,276 @@ placeholders with approved values. The sample does not load this JSON automatica "EPP_PROVIDER_CHANNEL": "sms", "EPP_PROVIDER_AUTH_MODE": "apiKey", "EPP_PROVIDER_ENDPOINT": "https:///", - "EPP_PROVIDER_TIMEOUT_MS": "1500", "KEY_VAULT_URL": "https://.vault.azure.net", - "AZURE_CLIENT_ID": "" + "AZURE_CLIENT_ID": "" } }, { - "id": "soprano-primary", + "id": "infobip-secondary", "settings": { - "EPP_PROVIDER_NAME": "soprano", - "EPP_PROVIDER_CHANNEL": "voice", - "EPP_PROVIDER_AUTH_MODE": "oauth", - "EPP_PROVIDER_ENDPOINT": "https:///", - "EPP_PROVIDER_TIMEOUT_MS": "1500", - "EPP_PROVIDER_TENANT_ID": "", - "EPP_PROVIDER_SCOPE": "api:///.default", - "EPP_OUTBOUND_CLIENT_ID": "", - "EPP_OUTBOUND_MI_CLIENT_ID": "" + "EPP_PROVIDER_NAME": "infobip", + "EPP_PROVIDER_CHANNEL": "sms", + "EPP_PROVIDER_AUTH_MODE": "apiKey", + "EPP_PROVIDER_ENDPOINT": "https://", + "EPP_PROVIDER_ACCOUNT_NAME": "", + "KEY_VAULT_URL": "https://.vault.azure.net", + "AZURE_CLIENT_ID": "" } } ] ``` -Store **references and identifiers, not secret values**. The -[Telesign adapter](../javascript/src/functions/providers/telesign.js) requests the secret names -`telesign-api-key` and `telesign-customer-id`. Separate vaults let two accounts use those same -names. Different names in one vault require your own per-context credential specification; -do not mutate the shared adapter's `credentialSpec`. - -The [Soprano adapter](../javascript/src/functions/providers/soprano.js) uses OAuth. Keep its -provider tenant, resource scope, outbound application and managed identity together as one -configuration. These settings do not grant consent or provider entitlement. +Your startup loader must reject duplicate/missing IDs, mismatched channels/authentication modes, +unapproved endpoints, missing credential references and invalid deadline policy. Validate both +contexts before enabling the endpoint, not only after the primary fails. -Keep inbound authentication and decryption configuration separate. Selecting an outbound provider -must not change the key used to decrypt the incoming request. +The [Telesign adapter](../javascript/src/functions/providers/telesign.js) requests +`telesign-api-key` and `telesign-customer-id`. The +[Infobip adapter](../javascript/src/functions/providers/infobip.js) requests `infobip-api-key` +and appends `/sms/3/messages` to its configured **base URL**. Reject an operation suffix, query +or fragment in that base URL; reject or normalize a trailing slash before freezing settings. +The suffix must appear once. Set an account/destination-approved sender; the adapter's `Verify` +default is not proof that this sender is valid for your account. -## 3. Create isolated, long-lived credential contexts +Store no credential values in the configuration. Attach each referenced user-assigned managed +identity to the Function App and grant the vault-reading identity scoped secret-read access, +such as Key Vault Secrets User under RBAC or the approved access-policy equivalent. If using +the system-assigned identity, grant that identity instead. Client-ID settings do not attach +identities or grant access. Complete provider account, recipient, sender and consent prerequisites; +see [provider onboarding](ONBOARDING.md#complete-provider-authentication). -Open [credentials.js](../javascript/src/functions/credentials.js). The exported -`credentialTokenService` singleton keeps **one selected cache**. Passing it a different account -or authentication mode does not switch that cache. +## 3. Give each context its own credential service and lifecycle -Use a new `CredentialTokenService` for each context instead. Build the contexts once per worker, -not once per request. Consolidate the imports below with the existing imports in your customized -`SendOtp.js`: `readConfig` and `selectProvider` are already imported, so do not paste duplicate -`const` declarations. Add `CredentialTokenService` to the existing credentials import and retain -required logging helpers such as `reportRefreshFailure`. Remove the `credentialTokenService` -singleton import only after replacing all its request and lifecycle uses. +Open [credentials.js](../javascript/src/functions/credentials.js). Its singleton caches the +first selected configuration; passing another account to it does not switch its cache. +Create a long-lived service for each account instead, with an immutable configuration: ```javascript +// Consolidate with existing imports in your customized SendOtp.js; do not duplicate consts. const { readConfig } = require('./config'); const { selectProvider } = require('./providers'); -const { CredentialTokenService, reportRefreshFailure } = require('./credentials'); +const { CredentialTokenService } = require('./credentials'); -// CUSTOMER_LOAD_AND_VALIDATE_CONFIG is your startup-only configuration loader. -// It must validate the complete list before any context is used. -const entries = CUSTOMER_LOAD_AND_VALIDATE_CONFIG(); +// CUSTOMER_* helpers are your implementation, not repository APIs. +const entries = CUSTOMER_LOAD_AND_VALIDATE_ACCOUNTS_AND_POLICY(); const contexts = new Map(); - for (const entry of entries) { - if (!entry.id || contexts.has(entry.id)) { - throw new Error('Missing or duplicate provider context ID'); - } const settings = Object.freeze({ ...entry.settings }); const config = Object.freeze(readConfig(settings)); const provider = selectProvider(config.providerName); - if (!provider || config.providerAuthMode !== provider.authenticationMode) { - throw new Error('Invalid provider context'); - } contexts.set(entry.id, Object.freeze({ - config, - provider, - credentials: new CredentialTokenService(), + config, provider, credentials: new CredentialTokenService(), })); } ``` -Your loader must also validate required credential settings, allowed channels, approved endpoint -allowlists, timeout values, a bounded context count, and that every routing rule identifies exactly -one matching context. The checks above are not a complete configuration validator. -[providers/index.js](../javascript/src/functions/providers/index.js) supplies the fixed adapter -lookup; unknown names return `null`. - -Freeze both the settings and config objects: `readConfig` retains the settings object as `env`. -Never reuse a context's service for a different vault, account, OAuth scope or identity. A provider -name alone is not enough to identify cached credentials. - -## 4. Select the context only after evaluation returns +The loader must finish all checks in step 2 before this loop runs. Do not share caches between +accounts, even if they use the same adapter. Keep expiry, acquisition bounds, refresh coalescing +and sanitized failure reporting. Logical isolation within a worker is not a tenant security boundary. -In [SendOtp.js](../javascript/src/functions/SendOtp.js), find `if (evaluation)`. Leave that branch, -the preceding envelope checks, [JWE decryption](../javascript/src/functions/jwe.js), and the -complete-delivery-context check in place. Keep Easy Auth enabled at the application boundary. - -The order must remain: - -```text -Authenticate -> validate envelope -> decrypt -> check delivery context - -> evaluation? Return the existing nonce response - -> otherwise select one provider context -``` - -Immediately after the existing evaluation early return, replace the block beginning -`const provider = selectProvider(config.providerName)` through the end of `if (!provider)` with -the following. Keep the `logContext = providerContext(...)` and `provider_selected` lines that -follow it. This fragment uses the handler's existing response helpers and request variables: +For this walkthrough, acquire credentials lazily. In `SendOtp.js`, remove the singleton's +`startProviderCredentialRefresh` function and its `app.hook.appStart(...)` registration. Replace +the termination function, retain its registration once, and update the **bottom export** together: ```javascript -const contextId = routeByChannel[payload.channelName]; -const selectedContext = contexts.get(contextId); -if (!selectedContext) { - fail('provider_selection', 'unknown_provider', 400); - return respond(400, { - error: 'provider_delivery_failed', correlationId, requestId, - }); +function stopProviderCredentialRefresh() { + for (const { credentials } of contexts.values()) credentials.close(); } -const { provider, config: providerConfig, credentials } = selectedContext; -``` - -The existing channel check below this block will reject a mismatched channel in step 5. -Use only the startup-owned `routeByChannel` map for selection. Evaluation must not depend on -that selection or acquire provider credentials. An invalid encrypted evaluation still fails the -existing checks; do not add a shortcut that returns a nonce before decryption. - -## 5. Use the selected context for one dispatch - -Continue in [SendOtp.js](../javascript/src/functions/SendOtp.js). Keep its channel, authentication -mode and endpoint checks, but read their outbound settings from `providerConfig`. Keep using the -original `const config = readConfig()` for `config.decryptionKeyPem` and `config.expectedKeyId`. -Do not redeclare `config` in the handler or replace its inbound settings with the selected context. -After the evaluation return, change the existing checks' `config.providerChannel`, -`config.providerAuthMode` and `config.providerEndpoint` references to the corresponding -`providerConfig` properties; leave their failure branches intact. -Inside the existing credential-resolution `try` block, replace the singleton call with: - -```javascript -credential = await credentials.getCredentials( - provider.credentialSpec, - providerConfig, -); -``` - -Keep the existing completeness checks for API-key identity/secret and OAuth access token, and -the `credential_unavailable` error path. For OAuth logging, pass `providerConfig` to -`credentialContext` too. - -Keep the handler's [OtpDelivery](../javascript/src/functions/delivery.js) construction so phone, -message, locale and correlation handling stay unchanged. In the existing **request-build** -`try` block, replace only the `provider.createRequest` call: - -```javascript -providerRequest = provider.createRequest({ - channel, - endpoint: providerConfig.providerEndpoint, - delivery, - credential, - env: providerConfig.env, -}); -``` - -In the separate **transport** `try` block, replace only the `sendProviderRequest` call. -Keep both blocks' existing catches so build failures and transport failures retain their -different error classifications: - -```javascript -transportResponse = await sendProviderRequest( - providerRequest, - parseProviderTimeout(providerConfig.providerTimeoutMs), - logContext, -); +// Keep this registration once, replacing the original lifecycle wiring. +app.hook.appTerminate(stopProviderCredentialRefresh); +// At the bottom of the module; do not leave the removed startup function here. +module.exports = { stopProviderCredentialRefresh }; ``` -The imports and error handling for `sendProviderRequest` and `parseProviderTimeout` already exist -in `SendOtp.js`. See [providerTransport.js](../javascript/src/functions/providerTransport.js) -for URL checks, manual redirects and timeout behavior. Preserve those controls and your approved -endpoint allowlist. Do not forward the inbound authorization header. - -Keep `provider.interpretResponse(transportResponse)` and the existing -[response mapping](../javascript/src/functions/providerResult.js). Preserve success/block/failure -statuses, the success nonce and correlation ID, and the handler's error responses. A provider -timeout, rejection or unknown response must not trigger another account or provider. -**One request gets one dispatch: no fan-out, automatic fallback or resend.** - -Retain the [logging helpers](../javascript/src/functions/logging.js) and their fixed failure -classifications. If you add a context ID to telemetry, use an approved non-secret identifier. -Never log OTP/phone data, request bodies, credentials, raw provider descriptions or exceptions. +Remove the singleton import only after replacing its request/lifecycle uses. Each service begins +periodic refresh when first used. An evaluation request must not start credential acquisition; +already-running refresh is independent. Rebuild contexts on controlled restart when configuration +changes; never repurpose a live service for another account. -## 6. Update startup, refresh and shutdown together +## 4. Add durable operation state before enabling fallback -Find `startProviderCredentialRefresh` and `stopProviderCredentialRefresh` in -[SendOtp.js](../javascript/src/functions/SendOtp.js). They currently use the singleton. Replace -that lifecycle wiring as part of your customization; leaving it behind could warm the wrong account. +The repository has **no durable attempt store or idempotency contract**. Implement these before +adding a second possible submission. Agree on a stable, authorized logical-operation key with +the caller/test contract. Do not assume a generated invocation ID, correlation GUID or nonce is +an idempotency key. Bind the key to the authorized caller, operation and immutable request/policy +identity so another request cannot reuse it to obtain a different delivery. -Decide explicitly whether to prewarm approved contexts at startup or acquire credentials on the -first live request. A service starts its own periodic refresh when first used. Prewarming performs -credential I/O, so do not do it in an offline test without fake SDK boundaries. Report acquisition -failures through the existing safe reporting mechanism; never borrow another context's credentials. +Use transactional/conditional writes in a durable shared store, not an in-memory map or lock. +Atomically claim an operation across concurrent invocations and worker restarts. Persist intent +**before** entering each provider transport call. For example, your state machine can use: -If you choose startup prewarming, replace **both existing function bodies** with the following -pattern. This assumes every context has passed your startup validation and is approved for -credential acquisition. Keep the existing `reportRefreshFailure` import; remove the singleton -import once no handler or lifecycle code references it. - -```javascript -async function startProviderCredentialRefresh() { - for (const { provider, config: providerConfig, credentials } of contexts.values()) { - try { - await credentials.getCredentials(provider.credentialSpec, providerConfig); - } catch { - // Refresh failures are reported by the service; report initialization failures here. - if (!credentials.current) reportRefreshFailure('configuration'); - } - } -} - -function stopProviderCredentialRefresh() { - for (const { credentials } of contexts.values()) { - credentials.close(); - } -} -``` - -Keep the existing `app.hook.appStart(startProviderCredentialRefresh)` and -`app.hook.appTerminate(stopProviderCredentialRefresh)` registrations once each; do not add duplicate -hooks. Build and validate `contexts` before startup runs. With prewarming, retain both functions -in the existing bottom export: +```text +NEW -> CLAIMED -> PRIMARY_INTENT -> ACCEPTED | TERMINAL | UNCERTAIN + | -> VERIFIED_NONACCEPTANCE + -> PRIMARY_PRE_DISPATCH_FAILURE -```javascript -module.exports = { startProviderCredentialRefresh, stopProviderCredentialRefresh }; +PRIMARY_PRE_DISPATCH_FAILURE or VERIFIED_NONACCEPTANCE + -> SECONDARY_RESERVED (atomic, policy-approved, once) + -> SECONDARY_INTENT -> ACCEPTED | TERMINAL | UNCERTAIN ``` -For **lazy acquisition**, remove the app-start registration and the -`startProviderCredentialRefresh` function together. Retain the context-closing termination -function and its hook, and replace the bottom export with: +All transitions require the expected state/version and valid ownership. A duplicate invocation +must not send; handle its response using the agreed operation contract. Never reset an intent +or uncertain dispatched attempt after a timeout, lease expiry or restart. An abandoned intent +may mean a message was sent even if no result was saved. Fail closed when the store is unavailable. +Do not let stale owners send after a claim is revoked or let a new owner reclaim a reserved attempt. + +Record operation state, safe provider context ID and per-attempt classifications separately from +request logs. Do not store/log bodies, OTPs or credentials as attempt evidence. Retention and key +reuse rules must cover the caller's replay window. A store guard cannot atomically commit an +external provider send: it is **not provider-side idempotency or true exactly-once delivery**. +Real SAS/native fallback outside this endpoint can still produce duplicates; coordinate that +behavior with the owner rather than promising the local guard prevents it. + +## 5. Integrate one guarded fallback into the handler + +In [SendOtp.js](../javascript/src/functions/SendOtp.js), preserve authentication at the platform +boundary, [envelope validation](../javascript/src/functions/entraPayload.js), +[decryption](../javascript/src/functions/jwe.js), completeness checks and the existing evaluation +early return. Keep inbound `config.decryptionKeyPem` and `config.expectedKeyId` unchanged. +Only live validated requests proceed to the operation claim and provider attempts. + +Replace the flow from `const provider = selectProvider(config.providerName)` through the +single-provider result handling with your guarded orchestration. Extract the existing credential, +request-build, transport and response blocks into an attempt helper that takes one context and +returns a structured category. Use `providerConfig` for outbound settings and that context's +`credentials.getCredentials(provider.credentialSpec, providerConfig)`. Keep +[OtpDelivery](../javascript/src/functions/delivery.js) construction and content handling intact. + +The existing credential catch immediately returns 502, and transport/result failures immediately +return errors. Your attempt helper must report the stage to the orchestrator instead of hiding +it in a catch that calls the secondary. Preserve their final HTTP classifications when fallback +is denied. This is **policy pseudocode**, not executable glue or a complete handler: ```javascript -module.exports = { stopProviderCredentialRefresh }; +// Runs after existing validation/decryption/evaluation handling. +const operation = await CUSTOMER_ATOMIC_CLAIM_AUTHORIZED_OPERATION(); +if (!operation.owned) return CUSTOMER_DUPLICATE_RESPONSE_WITHOUT_SENDING(); + +const primary = await CUSTOMER_ATTEMPT_ONCE( + operation, contexts.get(providerOrder.primary), aggregateDeadline); +if (primary.category === 'accepted') return CUSTOMER_EXISTING_SUCCESS_RESPONSE(); +if (!['primary_pre_dispatch_failure', 'verified_nonacceptance'] + .includes(primary.category)) return CUSTOMER_TERMINAL_OR_UNCERTAIN_RESPONSE(primary); + +const approval = CUSTOMER_CHECK_ACCOUNT_POLICY_AND_EVIDENCE(primary); +if (!approval.allowed || !CUSTOMER_HAS_SECONDARY_BUDGET(aggregateDeadline)) + return CUSTOMER_TERMINAL_OR_UNCERTAIN_RESPONSE(primary); +if (!await CUSTOMER_ATOMIC_RESERVE_SECONDARY(operation, approval)) + return CUSTOMER_DUPLICATE_RESPONSE_WITHOUT_SENDING(); + +const secondary = await CUSTOMER_ATTEMPT_ONCE( + operation, contexts.get(providerOrder.secondary), aggregateDeadline); +return CUSTOMER_FINAL_RESPONSE_WITHOUT_ANOTHER_ATTEMPT(secondary); ``` -Leaving the removed startup function in `module.exports` causes a `ReferenceError` when the module -loads, before any request can run. In either variant, no lifecycle code should continue to resolve -credentials from the original process-wide provider selection. - -Keep cache expiry, refresh coalescing and acquisition bounds. Use a controlled worker restart when -changing startup-loaded configuration, rebuilding each context and its cache. Do not -modify a live service's inputs. Evaluation requests must not start credential acquisition, but -already-running background refresh is independent of a request's evaluation branch. - -Separate contexts provide logical isolation, not a tenant security boundary. A compromised worker -may reach every identity assigned to it. Review least privilege before approving this design. - -## 7. Test your changes offline first - -Start from the existing -[adapter tests](../javascript/test/provider-flow.test.js), -[credential-cache tests](../javascript/test/credential-cache.test.js) and -[handler tests](../javascript/test/sendotp.test.js). -They show how to fake credential acquisition and provider transport. Use synthetic delivery data -and a locally generated encryption key; make unexpected network calls fail. - -**Before editing the handler**, run the existing tests as a baseline. Use Node.js 22 and restore -missing JavaScript dependencies from the existing lockfile first. From the repository root in -PowerShell: +Every `CUSTOMER_*` helper, result category, classifier, store and deadline here is new customer +code. `CUSTOMER_ATTEMPT_ONCE` must reserve/persist transport intent and recheck ownership and +remaining budget immediately before sending. Only the primary's approved pre-dispatch or proven +nonacceptance outcome can reserve a secondary; secondary failure is final. +Persist each outcome, including policy/budget denial, before the final response. If saving a +result fails after intent, leave that attempt non-reclaimable; do not send through the secondary. + +**Do not use `result.httpStatus >= 400` or `catch -> sendSecondary`.** +[providerResult.js](../javascript/src/functions/providerResult.js) exposes `Continue`, `Fail` +and `Block`, not a safe-to-fallback receipt. Preserve Block as terminal and recognized +accepted/`PENDING` as accepted. A customer classifier must require provider-specific, +documented nonacceptance for this exact attempt and account policy approval. It must not +reinterpret generic `Fail` or a broad status range as proof. + +[providerTransport.js](../javascript/src/functions/providerTransport.js) collapses fetch and +response-body failures into `provider_timeout`/`provider_network_error`. Those errors do **not** +expose whether the request was transmitted. Treat them as uncertain after intent, including +connection errors; do not guess they happened before send. A primary credential failure can +be pre-dispatch only when control flow and operation state prove transport was never entered, +the failure is isolated to that primary, and account policy still permits the secondary. + +Keep endpoint allowlists, HTTPS checks, manual redirects and per-provider response mapping. +Never forward inbound authorization. Preserve the endpoint success nonce/correlation response +only for evaluation or an accepted delivery result; terminal/uncertain outcomes get no success +nonce. Continue using [safe logging](../javascript/src/functions/logging.js), with an explicit +per-attempt provider ID/classification and one final request outcome. + +## 6. Enforce one aggregate deadline + +Define an end-to-end budget with the service/caller owner; this guide provides no SLA or magic +timeout value. Include validation, durable-store waits, primary credential acquisition/transport, +secondary credential acquisition/transport and final response work. An operation must not receive +a fresh budget when a duplicate invocation arrives. + +The current credential service has its own acquisition bound and the transport its own timeout. +`parseProviderTimeout` normalizes a single provider timeout, **not** an aggregate deadline. +Add explicit remaining-budget checks and cancellation propagation in your customer integration. +Do not use `Promise.race` to return while a send continues in the background. If ownership, +deadline or cancellation changes after dispatch, record uncertainty and never fall back. +Reserve enough time for the whole secondary attempt and finalization, rechecking before intent; +insufficient budget means no secondary even when nonacceptance is otherwise eligible. + +## 7. Validate offline, then deploy only after separate approval + +Before editing, run the existing JavaScript tests from the repository root. Restore missing +dependencies from the existing lockfile only when needed: ```powershell -# Only if dependencies are not installed. +# Only if dependencies are missing. npm ci --prefix .\javascript --ignore-scripts --no-audit --no-fund -``` - -Then run: - -```powershell node --test .\javascript\test\provider-flow.test.js ` .\javascript\test\credential-cache.test.js ` .\javascript\test\sendotp.test.js ``` -**After customization**, update the test fixtures and singleton/lifecycle mocks to use your -context services before rerunning and extending these tests. The original single-provider fixtures -are not automatically valid for your customized handler. - -They are a starting point, not coverage for your new router. Add tests for concurrent requests to -different channel bindings, including separate SMS/voice test accounts of the same adapter. -Do not introduce two active SMS choices to run these tests. Check the exact endpoint, -credential, message and correlation ID for each request. Test unknown/duplicate/ambiguous routes, -channel mismatch, expired or failing credentials, provider errors, timeouts, malformed responses, -shutdown and configuration replacement. Failures must cause no second dispatch. - -Check that valid evaluation returns before routing/credentials/transport and invalid evaluation -still fails. Cover voice and every adapter you intend to enable. The shared -[contract fixtures](../tests/fixtures/contract.json) help preserve endpoint response behavior. - -## 8. Review and approve live rollout separately - -Before deployment, review identity/vault permissions, OAuth consent, provider entitlement, -sender/channel approval, rate limits, secret rotation and operational ownership. Rotate any exposed -credentials first. Offline checks do not establish any of these prerequisites. - -Use a separately approved nonproduction deployment. Verify authentication, encryption and policy -readback before authorized evaluation. A real SAS trigger can involve surrounding fallback -behavior, so do not assume it is harmless because this Function's evaluation branch skips delivery. - -Only attempt live delivery with explicit authorization, an approved recipient, coordinated policy -and attempt limits, and a **one-attempt safety gate**. Provider acceptance is not proof of handset -delivery. Plan an explicit rollout and rollback; do not hide failures behind automatic fallback. -Nothing in this walkthrough authorizes a SAS trigger or provider send. - -### What needs deployment, configuration, or restart? - -Deploy your customized application code once, after its offline tests pass. Adding provider -settings to an unchanged checkout does not implement this guide: the shipped handler still -selects one provider per deployment. This customization hosts the SMS and voice account contexts -in **one Function App and one code package**, behind the existing `SendOtp` HTTP-triggered function. -It is not one deployment per provider, and it does not change the SAS-facing URL. - -| Change | Required action | -| --- | --- | -| Change the handler, channel map, an adapter, or a code-owned endpoint allowlist | Build, test and deploy a new application package. | -| Change account references, scopes or identities in the startup configuration loader | Validate the complete configuration, apply it, and perform a controlled worker restart. If configuration is embedded in code, deploy a new package instead. | -| Rotate a referenced secret | Update the approved secret and verify credential-cache refresh. Updating the vault alone is not proof that running workers use the new version; use a controlled restart when needed. Update pinned secret-version references explicitly. | -| Repeat an authorized test with unchanged code and configuration | No redeployment is needed. Verify the deployed package identifier, routing, authentication and credential readiness first. | -| Update only this documentation | No runtime deployment is needed. | - -The existing deployment automation does not create your custom account contexts, channel map or -configuration loader. Adapt your deployment process to provision only the required identities, -vault permissions and provider consent, then deploy the same reviewed package and validated -configuration to each intended instance. Preserve inbound authentication and decryption settings. -Treat code and configuration as one rollout, and retain their previous versions for rollback. - -### Replace the active SMS provider with Infobip - -This is an optional administrative replacement of Telesign for SMS, not an additional caller -choice. `/api/SendOtp` and the Soprano voice binding remain unchanged. - -1. Replace the `telesign-primary` entry in your custom loader's active configuration with the - following `infobip-primary` entry. Keep the Soprano entry. Supply approved account values, - never secret values: - -```json -{ - "id": "infobip-primary", - "settings": { - "EPP_PROVIDER_NAME": "infobip", - "EPP_PROVIDER_CHANNEL": "sms", - "EPP_PROVIDER_AUTH_MODE": "apiKey", - "EPP_PROVIDER_ENDPOINT": "https://", - "EPP_PROVIDER_ACCOUNT_NAME": "", - "EPP_PROVIDER_TIMEOUT_MS": "1500", - "KEY_VAULT_URL": "https://.vault.azure.net", - "AZURE_CLIENT_ID": "" - } -} -``` - -The [Infobip adapter](../javascript/src/functions/providers/infobip.js) reads `infobip-api-key` -and appends `/sms/3/messages` to the configured **base URL**. Do not put that operation suffix -in `EPP_PROVIDER_ENDPOINT`: it must be appended exactly once. In your startup loader, validate -the approved HTTPS base host with no operation path, query or fragment, and either reject a -trailing slash or remove it before freezing the settings; otherwise the adapter creates a double -separator. The adapter defaults to `Verify` if the sender setting is absent; that is not a -universally valid sender. Set the sender approved for your Infobip account and destination. - -2. Complete the account, sender, destination and vault-access prerequisites before activating - the replacement. Create a new long-lived credential service/cache for the Infobip context; - never reuse the Telesign cache with different inputs. - -Attach each referenced user-assigned managed identity to the Function App and grant the -vault-reading identity scoped secret-read access to its intended -vault, such as **Key Vault Secrets User** under RBAC or the approved access-policy equivalent. -If using the Function App's system-assigned identity, grant that identity instead. Context IDs and -managed-identity client-ID settings do not attach identities or grant permissions. See the -[provider credential onboarding guidance](ONBOARDING.md#complete-provider-authentication). -Keep shared inbound authentication/decryption separate from these outbound credentials. -The existing setup catalog does not provision this Infobip customization. - -3. Change only the `sms` value in `routeByChannel` from `telesign-primary` to `infobip-primary`. - Validate that each supported channel resolves to one configured context with the matching - channel/authentication mode and an approved provider endpoint. Reject missing or duplicate - configuration, and rerun the offline handler tests with the replacement binding. -4. The map in section 1 is code-owned, so build and deploy the reviewed package with the new map - and active configuration. For later changes to values already read by the startup loader, - validate and apply the complete configuration with a controlled worker restart. Close old - contexts and initialize fresh caches; do not switch an in-flight request or retry its delivery. - Verify the deployed package/configuration version before any separately authorized live check. - -SAS still sends to `/api/SendOtp`; it does not select the replacement provider. An SMS failure -must not return to Telesign, retry or send through both providers. [Front Door regional -failover](FRONTDOOR.md) remains separate: retain forwarding for `/api/SendOtp` and preserve -inbound authentication and origin/network restrictions. - -For an authorized live check, record unique correlations and durable intent-before-send attempt -guards. Stop on unexpected failures and do not retry uncertain sends automatically. A guard is -not provider idempotency or proof of exactly-once delivery. Record actual attempts, provider -acceptance and handset receipt separately: HTTP 200 or `PENDING` alone is not handset proof. -A sequential check does not establish concurrency, capacity, automatic failover, regional -resilience or real SAS integration. This replacement procedure authorizes no live operation. - -## Using .NET or Python instead - -**.NET:** Start with [SendOtp.cs](../dotnet/Functions/SendOtp.cs) and -[AppConfig.Read(IEnv)](../dotnet/Src/AppConfig.cs). Isolate the entire configuration, provider, -credential service and secret resolver per account. -[CredentialTokenService](../dotnet/Src/CredentialTokenService.cs) caches by `provider.Name`; -[SopranoProvider](../dotnet/Src/Providers/SopranoProvider.cs) retains its initial OAuth -identity/scope, and [SecretResolver](../dotnet/Src/SecretResolver.cs) retains its initial vault -client. Changing only the outer cache key is not enough. Update the singleton registrations and -lifecycle in [Program.cs](../dotnet/Program.cs), preserve -[PhoneProviderBase](../dotnet/Src/PhoneProviderBase.cs) transport/response behavior, and test -concurrent cold acquisition rather than assuming JavaScript's coalescing behavior. -Extend [SendOtpTests](../dotnet/tests/SendOtpTests.cs) and -[CredentialTokenServiceTests](../dotnet/tests/CredentialTokenServiceTests.cs). - -**Python:** Start with `_send_to_provider` and the evaluation branch in -[function_app.py](../python/function_app.py). Pass a selected context explicitly instead of -rereading process settings. Use [read_config(env)](../python/src/config.py), a separate +After customization, adapt the original singleton/lifecycle fixtures and extend the +[adapter](../javascript/test/provider-flow.test.js), +[credential](../javascript/test/credential-cache.test.js) and +[handler](../javascript/test/sendotp.test.js) tests. Use fake credentials, transport, clocks and +store faults; block external network. Verify: + +- Primary acceptance means zero secondary calls. Approved credential-before-transport failure + or explicitly modeled nonacceptance allows at most one secondary, only with guard and budget. +- Blocks, invalid input, generic provider errors, unknown/malformed responses and all uncertain + dispatches mean zero secondary calls. Secondary failure never restarts the sequence. +- Concurrent/retried logical operations and process restarts cannot reclaim intent/reservations. + Store failure, deadline exhaustion and cancellation fail closed. +- Evaluation returns before operation claims/provider calls; incomplete/decryption failures remain + errors. Credentials, endpoints and wire responses stay isolated between accounts. + +An offline policy-model experiment can check those branches, but a fake classifier does not +prove any real provider's nonacceptance semantics, and a fake store does not prove distributed +durability. Earlier routing/isolation tests are **not fallback tests**. No production +nonacceptance classifier is supplied here; keep response-based fallback disabled until reviewed. + +Use one Function App/package and the existing `/api/SendOtp` entry point. Deploy new code for the +orchestrator, classifier, durable-store integration or code-owned policy changes. Validate +startup-loaded configuration changes and perform a controlled restart. For secret rotation, +verify cache refresh; update pinned references explicitly. Repeating an unchanged approved test +or editing Markdown needs no redeployment. + +Preserve inbound authentication and origin/network restrictions, including any Front Door +forwarding. Regional failover is separate from provider fallback. Before live use, review +account permissions/consent, sender/recipient entitlement, durable-store guarantees and caller +retry/native-fallback behavior. Resolve exposed-credential rotation first. Any real SAS or +provider test needs separate authorization, an approved recipient, explicit attempt limits and +stop conditions for uncertainty. HTTP 200/`PENDING` is not handset receipt. Nothing here authorizes +a send or proves real SAS integration, capacity, regional resilience or live delivery. + +## .NET and Python integration pointers + +**.NET:** Start with [SendOtp.cs](../dotnet/Functions/SendOtp.cs), +[AppConfig.Read(IEnv)](../dotnet/Src/AppConfig.cs) and +[PhoneProviderBase](../dotnet/Src/PhoneProviderBase.cs). +[CredentialTokenService](../dotnet/Src/CredentialTokenService.cs) caches by provider name, while +[SopranoProvider](../dotnet/Src/Providers/SopranoProvider.cs) retains initialized OAuth state and +[SecretResolver](../dotnet/Src/SecretResolver.cs) retains its vault client. Isolate the full +account object graph and update [Program.cs](../dotnet/Program.cs) lifecycle registrations. +Review cancellation, concurrent acquisition and failure classification rather than assuming +the JavaScript behavior transfers unchanged. + +**Python:** Start with [function_app.py](../python/function_app.py), +[read_config(env)](../python/src/config.py) and [PhoneProviderBase](../python/src/provider.py). +Pass explicit account contexts instead of rereading process settings. Use a separate [CredentialTokenService](../python/src/credentials.py) and -[SecretResolver(env)](../python/src/secrets.py) per account; the module-level credential service -owns one selected cache. Update warmup/shutdown and preserve -[PhoneProviderBase](../python/src/provider.py) transport/response behavior. Extend -[test_engine.py](../python/tests/test_engine.py), -[test_credential_cache.py](../python/tests/test_credential_cache.py) and -[test_function_app.py](../python/tests/test_function_app.py). - -## Scope of this guidance - -An offline, in-memory adaptation of the registered JavaScript `SendOtp` handler checked this -SMS/Telesign and voice/Soprano design, same-provider account isolation, failures and evaluation -ordering. Its registered startup/shutdown callbacks and simulated refresh were also exercised. -Functions host registration, Azure SDK calls and provider HTTP were faked: this was not a deployed -or authenticated end-to-end test. Real OAuth/Key Vault, deployed authentication, production -refresh/scale-out and handset delivery still need your validation. .NET/Python pointers are -based on code inspection, not equivalent multi-context tests. - -Runtime adapters also include Infobip and Sinch; the [setup catalog](../setup/providers/catalog.json) -contains only Telesign and Soprano. Adapter availability does not imply setup coverage or account -entitlement. The original offline feasibility check performed no Azure, Graph, Key Vault, SAS or -provider operations. Report any subsequent live test of a customer customization separately, -including its deployed revision, authorized workload, provider responses and delivery limitations. -Such a test does not make the custom router part of the shipped sample. +[SecretResolver(env)](../python/src/secrets.py) per account; update warmup/shutdown. +Both runtimes still need customer-owned durable operation state, deadline enforcement and +reviewed fallback classification. These pointers are not evidence that fallback is implemented +or validated in either runtime. From 95c0251ac778a990cad225d987486da427de5abb Mon Sep 17 00:00:00 2001 From: Hou Chi Chan Date: Fri, 9 Oct 2026 14:26:05 -0700 Subject: [PATCH 14/14] Make public fallback guide provider-neutral Use role-based account IDs and supported-adapter placeholders instead of assigning named vendors primary or secondary status. Retain adapter-specific configuration guidance through neutral code links and preserve fallback safety rules. Co-authored-by: Copilot App <223556219+Copilot@users.noreply.github.com> Copilot-Session: 9314a109-37df-4021-8e75-bea995d57b23 --- docs/MULTI-PROVIDER-IMPLEMENTATION.md | 56 +++++++++++++++------------ 1 file changed, 31 insertions(+), 25 deletions(-) diff --git a/docs/MULTI-PROVIDER-IMPLEMENTATION.md b/docs/MULTI-PROVIDER-IMPLEMENTATION.md index 4a8f4a4..51e84bb 100644 --- a/docs/MULTI-PROVIDER-IMPLEMENTATION.md +++ b/docs/MULTI-PROVIDER-IMPLEMENTATION.md @@ -1,9 +1,10 @@ # Add conservative primary-to-secondary provider fallback You can adapt this application to try a secondary provider **only when the primary is known -not to have accepted the message**. Keep one SAS-facing URL, `/api/SendOtp`, and use Telesign -as the SMS primary and Infobip as the SMS secondary. Both serve the same channel; the caller -does not select an account or supply a provider header. +not to have accepted the message**. Keep one SAS-facing URL, `/api/SendOtp`, with a primary +account and a secondary account for the same SMS channel. You choose their order according to +your approved account policy; these roles do not imply a vendor ranking or recommendation. +The caller does not select an account or supply a provider header. The sample still ships with one provider per deployment and **no provider fallback**. This guide describes customer code you must implement and review. A timeout, lost response or generic @@ -17,8 +18,8 @@ Define the provider order in server-owned startup configuration: ```javascript const providerOrder = Object.freeze({ channel: 'sms', - primary: 'telesign-primary', - secondary: 'infobip-secondary', + primary: 'primary-account', + secondary: 'secondary-account', }); ``` @@ -37,7 +38,7 @@ Keep **routing order** separate from **fallback eligibility**. Use this policy: | Primary-only credential acquisition fails before transport is entered | Eligible only if this pre-dispatch fact is recorded, the account policy explicitly permits it, the durable operation guard allows the transition, and the request budget is sufficient. | | Provider-specific, documented proof that the primary did not accept the operation | Eligible only through a reviewed nonacceptance classifier, explicit account policy, the durable guard and sufficient budget. | -**Default to no fallback.** No real Telesign or Infobip response is designated safe to fall back +**Default to no fallback.** No real provider response is designated safe to fall back from by this guide. Obtain and review provider-specific semantics before enabling that path. An expired/rejected inbound token never reaches it. A provider credential failure must not bypass account suspension, consent requirements, entitlement restrictions or a provider block. @@ -46,31 +47,35 @@ account suspension, consent requirements, entitlement restrictions or a provider [config.js](../javascript/src/functions/config.js) accepts `readConfig(settings)`, so each account can use its own immutable settings without changing `process.env` per request. -The following is a **customer-owned configuration format**, not a file the sample loads: +The following is a **customer-owned configuration format**, not a file the sample loads. +The account IDs are role labels. Replace each `` with the adapter selected +for that account from the [existing lookup](../javascript/src/functions/providers/index.js); +placeholders are not valid runtime provider names. This example uses API-key authentication; +each account's settings must match its selected adapter's actual credential requirements. ```json [ { - "id": "telesign-primary", + "id": "primary-account", "settings": { - "EPP_PROVIDER_NAME": "telesign", + "EPP_PROVIDER_NAME": "", "EPP_PROVIDER_CHANNEL": "sms", "EPP_PROVIDER_AUTH_MODE": "apiKey", - "EPP_PROVIDER_ENDPOINT": "https:///", - "KEY_VAULT_URL": "https://.vault.azure.net", - "AZURE_CLIENT_ID": "" + "EPP_PROVIDER_ENDPOINT": "", + "KEY_VAULT_URL": "https://.vault.azure.net", + "AZURE_CLIENT_ID": "" } }, { - "id": "infobip-secondary", + "id": "secondary-account", "settings": { - "EPP_PROVIDER_NAME": "infobip", + "EPP_PROVIDER_NAME": "", "EPP_PROVIDER_CHANNEL": "sms", "EPP_PROVIDER_AUTH_MODE": "apiKey", - "EPP_PROVIDER_ENDPOINT": "https://", + "EPP_PROVIDER_ENDPOINT": "", "EPP_PROVIDER_ACCOUNT_NAME": "", - "KEY_VAULT_URL": "https://.vault.azure.net", - "AZURE_CLIENT_ID": "" + "KEY_VAULT_URL": "https://.vault.azure.net", + "AZURE_CLIENT_ID": "" } } ] @@ -80,13 +85,14 @@ Your startup loader must reject duplicate/missing IDs, mismatched channels/authe unapproved endpoints, missing credential references and invalid deadline policy. Validate both contexts before enabling the endpoint, not only after the primary fails. -The [Telesign adapter](../javascript/src/functions/providers/telesign.js) requests -`telesign-api-key` and `telesign-customer-id`. The -[Infobip adapter](../javascript/src/functions/providers/infobip.js) requests `infobip-api-key` -and appends `/sms/3/messages` to its configured **base URL**. Reject an operation suffix, query -or fragment in that base URL; reject or normalize a trailing slash before freezing settings. -The suffix must appear once. Set an account/destination-approved sender; the adapter's `Verify` -default is not proof that this sender is valid for your account. +For each selected [adapter](../javascript/src/functions/providers), inspect `credentialSpec` +for the required secret names and any account identifier, and `createRequest` for endpoint and +sender requirements. Some adapters accept a complete operation URL; others append an operation +path to a **base URL**. Use the approved HTTPS form expected by that adapter. For base URLs, +reject an already-appended operation suffix, query or fragment, and reject or normalize a trailing +slash before freezing settings so the operation path appears exactly once. +Set an account/destination-approved sender where required; an adapter's default sender is not +proof of approval. These requirements apply equally to primary and secondary roles. Store no credential values in the configuration. Attach each referenced user-assigned managed identity to the Function App and grant the vault-reading identity scoped secret-read access, @@ -316,7 +322,7 @@ a send or proves real SAS integration, capacity, regional resilience or live del [AppConfig.Read(IEnv)](../dotnet/Src/AppConfig.cs) and [PhoneProviderBase](../dotnet/Src/PhoneProviderBase.cs). [CredentialTokenService](../dotnet/Src/CredentialTokenService.cs) caches by provider name, while -[SopranoProvider](../dotnet/Src/Providers/SopranoProvider.cs) retains initialized OAuth state and +the [OAuth adapter implementation](../dotnet/Src/Providers) retains initialized OAuth state and [SecretResolver](../dotnet/Src/SecretResolver.cs) retains its vault client. Isolate the full account object graph and update [Program.cs](../dotnet/Program.cs) lifecycle registrations. Review cancellation, concurrent acquisition and failure classification rather than assuming