Skip to content

Azure provider cannot deploy with stock settings: five independent blockers #15046

Description

@rdreher

Environment

  • algo 2.0.0-beta, commit 20e22a8 (master)
  • Control machine: macOS, Python 3.12.14 (Homebrew), uv-managed .venv
  • ansible 12.3.0 / ansible-core 2.19.11rc1
  • azure.azcollection 3.12.0 (as bundled with ansible 12.3.0)
  • Azure CLI 2.75.0
  • Target image: config.cfg default, 0001-com-ubuntu-minimal-jammy-daily / minimal-22_04-daily-lts
  • Deployed successfully only after applying local workarounds for all five issues below

Summary

A stock clone with stock config.cfg cannot deploy to Azure. Five independent
problems block it, three of which produce error messages that point away from the
actual cause. I have local fixes for all of them and am happy to open PRs for the
ones where the approach isn't a design decision — see the end.


1. [azure] extra is incompatible with azure.azcollection

pyproject.toml pins:

azure = [
    "azure-mgmt-compute>=38.2.0",
    "azure-mgmt-network>=31.0.1",
    "azure-mgmt-resource>=26.0.0",
    ...
]

azure-mgmt-resource 24.0.0 split subscriptions, locks, and other
subpackages out of the distribution. In 26.0.0, azure/mgmt/resource/ contains
only resources/. But azure.azcollection 3.12.0 pins
azure-mgmt-resource==23.2.0 and azure_rm_common.py imports:

from azure.mgmt.resource.subscriptions import SubscriptionClient
from azure.mgmt.resource.locks import ManagementLockClient

Result:

Failed to import the required Python library (azure.mgmt.resource.subscriptions).

This is not fixable by hand, because playbooks/cloud-pre.yml runs
uv pip install '.[{{ cloud_provider_extra }}]' on every deploy — the
incompatible version is reinstalled immediately before the task that needs the
older layout.

Workaround: pin the extra to the collection's versions
(azure-mgmt-resource==23.2.0, azure-mgmt-compute==33.0.0,
azure-mgmt-network==28.0.0).

Note: the right fix may instead be to bump the collection, or to drop the
[azure] extra and defer to the collection's own requirements.txt — see #2.
I don't want to presume which.

2. [azure] extra covers 5 of the ~40 packages the collection imports

azure_rm_common.py imports roughly 40 SDK packages in a single try block,
so one missing import sets HAS_AZURE = False and disables every azure_rm_*
module. The reported error names neither the missing package nor the real
problem:

Failed to import the required Python library (ansible[azure] (azure >= 2.0.0)).
Please install them by running: pip install -r requirements.txt

(ansible[azure] as an extra hasn't existed since Ansible 2.10, which sends
users down a dead end.)

Algo's extra provides 5 of those packages. Deploying requires separately
installing the collection's own requirements:

ansible-galaxy collection install azure.azcollection --force
pip install -r .../ansible_collections/azure/azcollection/requirements.txt

Expect a pip dependency-resolver warning afterwards: azure-cli-core 2.75.0
pins msal==1.33.0b1, which requires cryptography<48, while algo requires
cryptography>=50.0.0. This appears benign — algo's only direct use is x25519
serialization in library/x25519_pubkey.py — but the two pins are formally
unsatisfiable together.

Suggestion: either document this install step for the Azure provider, or
have cloud-pre.yml install the collection's requirements.txt for
algo_provider == 'azure'.

3. No fallback to the Azure CLI's default subscription

docs/cloud-azure.md states that after az login, "you are able to deploy an
AlgoVPN instance without hassle." The role does not honor the CLI's default
subscription. roles/cloud-azure/tasks/prompts.yml:

subscription_id: "{{ azure_subscription_id | default(lookup('env', 'AZURE_SUBSCRIPTION_ID'), true) }}"

With neither set, an empty string is passed to azure_rm_deployment, the request
URL becomes /subscriptions//resourcegroups/..., and ARM parses the next path
segment as the subscription ID:

(InvalidSubscriptionId) The provided subscription identifier 'resourcegroups'
is malformed or invalid.

Suggested fix: when both are empty, fall back to
az account show --query id -o tsv, and fail with a message naming all three
options rather than proceeding with an empty value. Five lines; happy to PR.

4. Region list is a stale hardcoded snapshot, with an off-by-one default

roles/cloud-azure/defaults/main.yml hardcodes 88 regions from a one-time
az account list-locations run. Two problems:

26 entries cannot host a VM. asia, asiapacific, australia, brazil,
canada, europe, france, germany, global, india, israel, italy,
japan, korea, newzealand, norway, poland, qatar, singapore,
southafrica, sweden, switzerland, uae, uk, unitedstates,
unitedstateseuap are logical geo groupings (metadata.regionType == "Logical").
A further 11 are *stage regions. Selecting any of them from the menu fails.

The default is computed unsafely. prompts.yml:

default_region: >-
  {% for r in azure_regions %}{%- if r['name'] == "eastus" %}{{ loop.index }}{% endif %}{%- endfor %}

If eastus is absent from the list, this yields an empty string; | int turns
it into 0; and azure_regions[0 - 1] in main.yml silently selects the
last region in the list. The user gets a deployment in a region they did not
choose, with no warning.

Suggested fix: the off-by-one is a clear bug and worth fixing on its own
(fall back to index 1, or fail explicitly). Filtering the list to
regionType == "Physical" would also help. Note the menu cannot be filtered to
what a given subscription is entitled to without querying at runtime, which is
how a user ends up choosing a region where they have no capacity — see #5's
context.

5. Privacy role requires rsyslog, which algo's default Azure image lacks

config.cfg defaults to 0001-com-ubuntu-minimal-jammy-daily for Azure, which
ships without rsyslog and logs only to systemd-journald. The privacy role writes
/etc/rsyslog.d/*.conf, validates with rsyslogd -N1, and notifies a
restart rsyslog handler. All three fail:

TASK [privacy : Test rsyslog configuration]
Error executing command: [Errno 2] No such file or directory: b'rsyslogd'

and after gating that task:

TASK [privacy : restart rsyslog]
Could not find the requested service rsyslog: host

Notifiers live in log_filtering.yml (4), log_rotation.yml (3), and
advanced_privacy.yml (1).

Suggested fix: stat /usr/sbin/rsyslogd once in
roles/privacy/tasks/main.yml, register a fact, and gate the rsyslog-specific
tasks and the handler on it. Skipping is correct rather than a privacy
regression: advanced_privacy.yml already sets Storage=volatile in
journald.conf and removes /var/log/journal, so on a minimal image there is
no persistent log to filter. Happy to PR.


Knock-on effect: a partially-applied run leaves a silently broken server

Worth recording because it cost the most time to diagnose. When the run aborts
at #5, the VM is already provisioned and server.yml has run the strongswan
role, so client profiles are written (they're generated locally) and everything
looks successful. Two things are missing, both downstream of the abort:

  1. restart strongswan never fires. Ansible discards queued handlers when a
    play fails, so charon keeps running with the config it had at startup.
    ipsec statusall reports a healthy daemon with an empty Connections: block.
  2. distribute_keys.yml never copies the CA certificate to
    /etc/ipsec.d/cacerts/ca.crt. The server cert and key are present, so
    Connections: looks correct once loaded, but chain validation fails:
received 1 cert requests for an unknown ca
received end entity cert "CN=<user>"
no issuer certificate found for "CN=<user>"
  issuer is "CN=<server ip>"
no trusted ECDSA public key found for '<user>@<uuid>.algo'
generating IKE_AUTH response 1 [ N(AUTH_FAILED) ]

On macOS this surfaces only as a VPN toggle that flips back with no error
message. Diagnosis also requires overriding strongswan_log_level: -1, which
silences charon by default.

Two notes for anyone recovering such a server: ipsec rereadall does not
reliably load a newly added CA certificate (ipsec listcacerts stays empty) —
a full ipsec restart is needed. And /var/log/syslog doesn't exist on the
minimal image, so charon's --use-syslog output is only reachable via
journalctl -t charon.

This isn't a separate bug so much as an argument for #5 being higher severity
than it looks: the failure mode isn't "playbook stops," it's "playbook stops and
leaves a server that appears deployed and cannot authenticate."


Context on VM size, not a bug

For completeness, since it interacts with #4: config.cfg defaults Azure to
Standard_B1S, a first-generation burstable SKU with frequent capacity
restrictions in smaller regions. Deployments to South India and West India both
failed with:

SkuNotAvailable: Following SKUs have failed for Capacity Restrictions:
Standard_B1S ... is currently not available in location 'WestIndia'

Standard_B2ats_v2 (2 vCPU, 1 GiB, same memory, similar price) deployed without
issue. A note in docs/cloud-azure.md might save people some guessing, since
az vm list-skus does not report physical capacity restrictions and so cannot
be checked in advance.


Offer

I have working fixes for all of the above. I'd propose PRs for #3, the
off-by-one half of #4, and #5, since those don't involve a design decision.
For #1 and #2 I'd rather hear which direction you prefer before writing
anything — pinning algo's extra to the collection, bumping the collection, or
dropping the extra in favor of the collection's own requirements file.

Fixes have been verified on one successful end-to-end deployment
(Central India, Standard_B2ats_v2, IKEv2 + WireGuard both working); they have
not been run through the repo's CI.

No activity

Activity on this issue will appear here.

Activity

Sign up for free to join this conversation on GitHub. Already have an account? Sign in to comment

Metadata

Metadata

Assignees

No one assigned

    Labels

    No labels
    No labels

    Type

    No type

    Projects

    No projects

      Milestone

      No milestone

      Relationships

      None yet

      Development

      No branches or pull requests

      Issue actions