Environment
- algo
2.0.0-beta, commit 20e22a8 (master)
- Control machine: macOS, Python 3.12.14 (Homebrew), uv-managed
.venv
ansible 12.3.0 / ansible-core 2.19.11rc1
azure.azcollection 3.12.0 (as bundled with ansible 12.3.0)
- Azure CLI 2.75.0
- Target image:
config.cfg default, 0001-com-ubuntu-minimal-jammy-daily / minimal-22_04-daily-lts
- Deployed successfully only after applying local workarounds for all five issues below
Summary
A stock clone with stock config.cfg cannot deploy to Azure. Five independent
problems block it, three of which produce error messages that point away from the
actual cause. I have local fixes for all of them and am happy to open PRs for the
ones where the approach isn't a design decision — see the end.
1. [azure] extra is incompatible with azure.azcollection
pyproject.toml pins:
azure = [
"azure-mgmt-compute>=38.2.0",
"azure-mgmt-network>=31.0.1",
"azure-mgmt-resource>=26.0.0",
...
]
azure-mgmt-resource 24.0.0 split subscriptions, locks, and other
subpackages out of the distribution. In 26.0.0, azure/mgmt/resource/ contains
only resources/. But azure.azcollection 3.12.0 pins
azure-mgmt-resource==23.2.0 and azure_rm_common.py imports:
from azure.mgmt.resource.subscriptions import SubscriptionClient
from azure.mgmt.resource.locks import ManagementLockClient
Result:
Failed to import the required Python library (azure.mgmt.resource.subscriptions).
This is not fixable by hand, because playbooks/cloud-pre.yml runs
uv pip install '.[{{ cloud_provider_extra }}]' on every deploy — the
incompatible version is reinstalled immediately before the task that needs the
older layout.
Workaround: pin the extra to the collection's versions
(azure-mgmt-resource==23.2.0, azure-mgmt-compute==33.0.0,
azure-mgmt-network==28.0.0).
Note: the right fix may instead be to bump the collection, or to drop the
[azure] extra and defer to the collection's own requirements.txt — see #2.
I don't want to presume which.
2. [azure] extra covers 5 of the ~40 packages the collection imports
azure_rm_common.py imports roughly 40 SDK packages in a single try block,
so one missing import sets HAS_AZURE = False and disables every azure_rm_*
module. The reported error names neither the missing package nor the real
problem:
Failed to import the required Python library (ansible[azure] (azure >= 2.0.0)).
Please install them by running: pip install -r requirements.txt
(ansible[azure] as an extra hasn't existed since Ansible 2.10, which sends
users down a dead end.)
Algo's extra provides 5 of those packages. Deploying requires separately
installing the collection's own requirements:
ansible-galaxy collection install azure.azcollection --force
pip install -r .../ansible_collections/azure/azcollection/requirements.txt
Expect a pip dependency-resolver warning afterwards: azure-cli-core 2.75.0
pins msal==1.33.0b1, which requires cryptography<48, while algo requires
cryptography>=50.0.0. This appears benign — algo's only direct use is x25519
serialization in library/x25519_pubkey.py — but the two pins are formally
unsatisfiable together.
Suggestion: either document this install step for the Azure provider, or
have cloud-pre.yml install the collection's requirements.txt for
algo_provider == 'azure'.
3. No fallback to the Azure CLI's default subscription
docs/cloud-azure.md states that after az login, "you are able to deploy an
AlgoVPN instance without hassle." The role does not honor the CLI's default
subscription. roles/cloud-azure/tasks/prompts.yml:
subscription_id: "{{ azure_subscription_id | default(lookup('env', 'AZURE_SUBSCRIPTION_ID'), true) }}"
With neither set, an empty string is passed to azure_rm_deployment, the request
URL becomes /subscriptions//resourcegroups/..., and ARM parses the next path
segment as the subscription ID:
(InvalidSubscriptionId) The provided subscription identifier 'resourcegroups'
is malformed or invalid.
Suggested fix: when both are empty, fall back to
az account show --query id -o tsv, and fail with a message naming all three
options rather than proceeding with an empty value. Five lines; happy to PR.
4. Region list is a stale hardcoded snapshot, with an off-by-one default
roles/cloud-azure/defaults/main.yml hardcodes 88 regions from a one-time
az account list-locations run. Two problems:
26 entries cannot host a VM. asia, asiapacific, australia, brazil,
canada, europe, france, germany, global, india, israel, italy,
japan, korea, newzealand, norway, poland, qatar, singapore,
southafrica, sweden, switzerland, uae, uk, unitedstates,
unitedstateseuap are logical geo groupings (metadata.regionType == "Logical").
A further 11 are *stage regions. Selecting any of them from the menu fails.
The default is computed unsafely. prompts.yml:
default_region: >-
{% for r in azure_regions %}{%- if r['name'] == "eastus" %}{{ loop.index }}{% endif %}{%- endfor %}
If eastus is absent from the list, this yields an empty string; | int turns
it into 0; and azure_regions[0 - 1] in main.yml silently selects the
last region in the list. The user gets a deployment in a region they did not
choose, with no warning.
Suggested fix: the off-by-one is a clear bug and worth fixing on its own
(fall back to index 1, or fail explicitly). Filtering the list to
regionType == "Physical" would also help. Note the menu cannot be filtered to
what a given subscription is entitled to without querying at runtime, which is
how a user ends up choosing a region where they have no capacity — see #5's
context.
5. Privacy role requires rsyslog, which algo's default Azure image lacks
config.cfg defaults to 0001-com-ubuntu-minimal-jammy-daily for Azure, which
ships without rsyslog and logs only to systemd-journald. The privacy role writes
/etc/rsyslog.d/*.conf, validates with rsyslogd -N1, and notifies a
restart rsyslog handler. All three fail:
TASK [privacy : Test rsyslog configuration]
Error executing command: [Errno 2] No such file or directory: b'rsyslogd'
and after gating that task:
TASK [privacy : restart rsyslog]
Could not find the requested service rsyslog: host
Notifiers live in log_filtering.yml (4), log_rotation.yml (3), and
advanced_privacy.yml (1).
Suggested fix: stat /usr/sbin/rsyslogd once in
roles/privacy/tasks/main.yml, register a fact, and gate the rsyslog-specific
tasks and the handler on it. Skipping is correct rather than a privacy
regression: advanced_privacy.yml already sets Storage=volatile in
journald.conf and removes /var/log/journal, so on a minimal image there is
no persistent log to filter. Happy to PR.
Knock-on effect: a partially-applied run leaves a silently broken server
Worth recording because it cost the most time to diagnose. When the run aborts
at #5, the VM is already provisioned and server.yml has run the strongswan
role, so client profiles are written (they're generated locally) and everything
looks successful. Two things are missing, both downstream of the abort:
restart strongswan never fires. Ansible discards queued handlers when a
play fails, so charon keeps running with the config it had at startup.
ipsec statusall reports a healthy daemon with an empty Connections: block.
distribute_keys.yml never copies the CA certificate to
/etc/ipsec.d/cacerts/ca.crt. The server cert and key are present, so
Connections: looks correct once loaded, but chain validation fails:
received 1 cert requests for an unknown ca
received end entity cert "CN=<user>"
no issuer certificate found for "CN=<user>"
issuer is "CN=<server ip>"
no trusted ECDSA public key found for '<user>@<uuid>.algo'
generating IKE_AUTH response 1 [ N(AUTH_FAILED) ]
On macOS this surfaces only as a VPN toggle that flips back with no error
message. Diagnosis also requires overriding strongswan_log_level: -1, which
silences charon by default.
Two notes for anyone recovering such a server: ipsec rereadall does not
reliably load a newly added CA certificate (ipsec listcacerts stays empty) —
a full ipsec restart is needed. And /var/log/syslog doesn't exist on the
minimal image, so charon's --use-syslog output is only reachable via
journalctl -t charon.
This isn't a separate bug so much as an argument for #5 being higher severity
than it looks: the failure mode isn't "playbook stops," it's "playbook stops and
leaves a server that appears deployed and cannot authenticate."
Context on VM size, not a bug
For completeness, since it interacts with #4: config.cfg defaults Azure to
Standard_B1S, a first-generation burstable SKU with frequent capacity
restrictions in smaller regions. Deployments to South India and West India both
failed with:
SkuNotAvailable: Following SKUs have failed for Capacity Restrictions:
Standard_B1S ... is currently not available in location 'WestIndia'
Standard_B2ats_v2 (2 vCPU, 1 GiB, same memory, similar price) deployed without
issue. A note in docs/cloud-azure.md might save people some guessing, since
az vm list-skus does not report physical capacity restrictions and so cannot
be checked in advance.
Offer
I have working fixes for all of the above. I'd propose PRs for #3, the
off-by-one half of #4, and #5, since those don't involve a design decision.
For #1 and #2 I'd rather hear which direction you prefer before writing
anything — pinning algo's extra to the collection, bumping the collection, or
dropping the extra in favor of the collection's own requirements file.
Fixes have been verified on one successful end-to-end deployment
(Central India, Standard_B2ats_v2, IKEv2 + WireGuard both working); they have
not been run through the repo's CI.
Environment
2.0.0-beta, commit20e22a8(master).venvansible12.3.0 /ansible-core2.19.11rc1azure.azcollection3.12.0 (as bundled withansible12.3.0)config.cfgdefault,0001-com-ubuntu-minimal-jammy-daily/minimal-22_04-daily-ltsSummary
A stock clone with stock
config.cfgcannot deploy to Azure. Five independentproblems block it, three of which produce error messages that point away from the
actual cause. I have local fixes for all of them and am happy to open PRs for the
ones where the approach isn't a design decision — see the end.
1.
[azure]extra is incompatible withazure.azcollectionpyproject.tomlpins:azure-mgmt-resource24.0.0 splitsubscriptions,locks, and othersubpackages out of the distribution. In 26.0.0,
azure/mgmt/resource/containsonly
resources/. Butazure.azcollection3.12.0 pinsazure-mgmt-resource==23.2.0andazure_rm_common.pyimports:Result:
This is not fixable by hand, because
playbooks/cloud-pre.ymlrunsuv pip install '.[{{ cloud_provider_extra }}]'on every deploy — theincompatible version is reinstalled immediately before the task that needs the
older layout.
Workaround: pin the extra to the collection's versions
(
azure-mgmt-resource==23.2.0,azure-mgmt-compute==33.0.0,azure-mgmt-network==28.0.0).Note: the right fix may instead be to bump the collection, or to drop the
[azure]extra and defer to the collection's ownrequirements.txt— see #2.I don't want to presume which.
2.
[azure]extra covers 5 of the ~40 packages the collection importsazure_rm_common.pyimports roughly 40 SDK packages in a singletryblock,so one missing import sets
HAS_AZURE = Falseand disables everyazure_rm_*module. The reported error names neither the missing package nor the real
problem:
(
ansible[azure]as an extra hasn't existed since Ansible 2.10, which sendsusers down a dead end.)
Algo's extra provides 5 of those packages. Deploying requires separately
installing the collection's own requirements:
Expect a
pipdependency-resolver warning afterwards:azure-cli-core 2.75.0pins
msal==1.33.0b1, which requirescryptography<48, while algo requirescryptography>=50.0.0. This appears benign — algo's only direct use is x25519serialization in
library/x25519_pubkey.py— but the two pins are formallyunsatisfiable together.
Suggestion: either document this install step for the Azure provider, or
have
cloud-pre.ymlinstall the collection'srequirements.txtforalgo_provider == 'azure'.3. No fallback to the Azure CLI's default subscription
docs/cloud-azure.mdstates that afteraz login, "you are able to deploy anAlgoVPN instance without hassle." The role does not honor the CLI's default
subscription.
roles/cloud-azure/tasks/prompts.yml:With neither set, an empty string is passed to
azure_rm_deployment, the requestURL becomes
/subscriptions//resourcegroups/..., and ARM parses the next pathsegment as the subscription ID:
Suggested fix: when both are empty, fall back to
az account show --query id -o tsv, and fail with a message naming all threeoptions rather than proceeding with an empty value. Five lines; happy to PR.
4. Region list is a stale hardcoded snapshot, with an off-by-one default
roles/cloud-azure/defaults/main.ymlhardcodes 88 regions from a one-timeaz account list-locationsrun. Two problems:26 entries cannot host a VM.
asia,asiapacific,australia,brazil,canada,europe,france,germany,global,india,israel,italy,japan,korea,newzealand,norway,poland,qatar,singapore,southafrica,sweden,switzerland,uae,uk,unitedstates,unitedstateseuapare logical geo groupings (metadata.regionType == "Logical").A further 11 are
*stageregions. Selecting any of them from the menu fails.The default is computed unsafely.
prompts.yml:If
eastusis absent from the list, this yields an empty string;| intturnsit into
0; andazure_regions[0 - 1]inmain.ymlsilently selects thelast region in the list. The user gets a deployment in a region they did not
choose, with no warning.
Suggested fix: the off-by-one is a clear bug and worth fixing on its own
(fall back to index 1, or fail explicitly). Filtering the list to
regionType == "Physical"would also help. Note the menu cannot be filtered towhat a given subscription is entitled to without querying at runtime, which is
how a user ends up choosing a region where they have no capacity — see #5's
context.
5. Privacy role requires rsyslog, which algo's default Azure image lacks
config.cfgdefaults to0001-com-ubuntu-minimal-jammy-dailyfor Azure, whichships without rsyslog and logs only to systemd-journald. The privacy role writes
/etc/rsyslog.d/*.conf, validates withrsyslogd -N1, and notifies arestart rsysloghandler. All three fail:and after gating that task:
Notifiers live in
log_filtering.yml(4),log_rotation.yml(3), andadvanced_privacy.yml(1).Suggested fix: stat
/usr/sbin/rsyslogdonce inroles/privacy/tasks/main.yml, register a fact, and gate the rsyslog-specifictasks and the handler on it. Skipping is correct rather than a privacy
regression:
advanced_privacy.ymlalready setsStorage=volatileinjournald.confand removes/var/log/journal, so on a minimal image there isno persistent log to filter. Happy to PR.
Knock-on effect: a partially-applied run leaves a silently broken server
Worth recording because it cost the most time to diagnose. When the run aborts
at #5, the VM is already provisioned and
server.ymlhas run the strongswanrole, so client profiles are written (they're generated locally) and everything
looks successful. Two things are missing, both downstream of the abort:
restart strongswannever fires. Ansible discards queued handlers when aplay fails, so charon keeps running with the config it had at startup.
ipsec statusallreports a healthy daemon with an emptyConnections:block.distribute_keys.ymlnever copies the CA certificate to/etc/ipsec.d/cacerts/ca.crt. The server cert and key are present, soConnections:looks correct once loaded, but chain validation fails:On macOS this surfaces only as a VPN toggle that flips back with no error
message. Diagnosis also requires overriding
strongswan_log_level: -1, whichsilences charon by default.
Two notes for anyone recovering such a server:
ipsec rereadalldoes notreliably load a newly added CA certificate (
ipsec listcacertsstays empty) —a full
ipsec restartis needed. And/var/log/syslogdoesn't exist on theminimal image, so charon's
--use-syslogoutput is only reachable viajournalctl -t charon.This isn't a separate bug so much as an argument for #5 being higher severity
than it looks: the failure mode isn't "playbook stops," it's "playbook stops and
leaves a server that appears deployed and cannot authenticate."
Context on VM size, not a bug
For completeness, since it interacts with #4:
config.cfgdefaults Azure toStandard_B1S, a first-generation burstable SKU with frequent capacityrestrictions in smaller regions. Deployments to South India and West India both
failed with:
Standard_B2ats_v2(2 vCPU, 1 GiB, same memory, similar price) deployed withoutissue. A note in
docs/cloud-azure.mdmight save people some guessing, sinceaz vm list-skusdoes not report physical capacity restrictions and so cannotbe checked in advance.
Offer
I have working fixes for all of the above. I'd propose PRs for #3, the
off-by-one half of #4, and #5, since those don't involve a design decision.
For #1 and #2 I'd rather hear which direction you prefer before writing
anything — pinning algo's extra to the collection, bumping the collection, or
dropping the extra in favor of the collection's own requirements file.
Fixes have been verified on one successful end-to-end deployment
(Central India,
Standard_B2ats_v2, IKEv2 + WireGuard both working); they havenot been run through the repo's CI.