Skip to content

Instance metadata and user data

Read this before writing cloud-init for this platform. One deliberate difference from AWS changes how configuration reaches your machine, and it is not visible from the API side.

Configuration arrives on a config drive, not over the network

Section titled “Configuration arrives on a config drive, not over the network”
force_config_drive = true

Every guest on this platform is handed its configuration on a config drive — a small read-only volume attached to the machine, labelled config-2 — and never over the metadata network.

the compute service normally attaches a config drive only when an instance asks for one, and the EC2 API has no parameter for asking. So no consumer of this cloud could ask; forcing it makes it the platform’s behaviour rather than a per-launch accident.

The reason is a dataplane one. A guest reaches 169.254.169.254 through the router of its own network — and a machine that has not been configured yet is exactly the machine least able to fix a fault on that path. The config drive is the only delivery that does not depend on the network coming up first.

For Talos this matters twice over: it looks for the config-2 volume first and falls back to the metadata address, so with this set it never touches the network to install itself.

What you should take from it: if your machine did not get its keys or its user data, the network is not the first thing to check. Mount the drive and look.

Terminal window
blkid -t LABEL=config-2
mkdir -p /mnt/cfg && mount /dev/disk/by-label/config-2 /mnt/cfg
cat /mnt/cfg/openstack/latest/user_data
cat /mnt/cfg/openstack/latest/meta_data.json

There is one — ec2-api-metadata runs alongside the EC2 API — and it answers on the usual link-local address from inside an instance:

http://169.254.169.254/latest/meta-data/
http://169.254.169.254/latest/user-data

It is the EC2-shaped metadata service, so the familiar paths work: instance-id, instance-type, local-ipv4, public-ipv4, hostname, placement/availability-zone, public-keys/0/openssh-key.

It is a fallback and a convenience here, not the delivery mechanism. Everything it serves is also on the config drive.

AWSHere
IMDSv2 session tokens (PUT /latest/api/token, X-aws-ec2-metadata-token)Not supported. the compute service’s metadata service has no session-token mode
HttpTokens=requiredRefused, not silently downgraded — UnsupportedOperation
Instance tags in metadata (InstanceMetadataTags)disabled
IPv6 metadata endpointdisabled
IAM role credentials at iam/security-credentials/Nothing to serve. There are no instance roles — see Identity, as deployed

The layer’s defaults, which are the only values MetadataOptions will accept:

HttpTokens=optional HttpEndpoint=enabled HttpPutResponseHopLimit=1
HttpProtocolIpv6=disabled InstanceMetadataTags=disabled

RunInstances sends MetadataOptions only when Terraform has a metadata_options block; any member that differs from the above is rejected rather than quietly ignored.

[!warning]

The metadata endpoint is unauthenticated from inside the guest, and HttpTokens=required cannot be set. IMDSv2 exists on AWS specifically to blunt SSRF — an application tricked into fetching a URL cannot reach IMDS because it will not send the token header. That defence is not available here.

There are no role credentials to steal, which is the worst of what SSRF gets on AWS. But your SSH public key, hostname and network layout are readable by anything that can make an HTTP request from inside the machine.

ParameterUserData on RunInstances
EncodingBase64
Limit16384 bytes decoded — AWS’s limit, enforced, even though the compute service would allow 65535
LoggingNever logged. The model marks the shape sensitive
DeliveryConfig drive, and the metadata service
Terminal window
aws --endpoint-url https://ec2.shelfcs.com --region hel1 \
ec2 run-instances --image-id ami-… --instance-type cd-standard-2-4 \
--key-name mykey --user-data file://cloud-init.yaml
resource "aws_instance" "app" {
ami = data.aws_ami.debian.id
instance_type = "cd-standard-2-4"
key_name = aws_key_pair.mine.key_name
user_data = file("cloud-init.yaml")
}

Over 16384 bytes decoded is InvalidParameterValue. If your configuration is bigger than that, put a fetch in the user data and the payload somewhere your machine can reach — noting that there is no object storage on this platform to put it in.

Login userSet by the image, never by us — debian, ubuntu, rocky, almalinux, fedora, opensuse, arch. See Machine images
Root SSHDisabled by the vendor image. Your key goes to the login user
AddressOn 10.100.0.0/24, gateway and NAT at 10.100.0.1
DNSQuad9 — 9.9.9.9, 149.112.112.112
Root disk20–160 GB by shape, deleted with the instance

Talos takes no cloud-init: it reads Talos machine configuration from the same user-data channel, and has no SSH or login user at all.

[!warning]

cloud-init’s runcmd swallows failures. Every command in a runcmd block can fail and cloud-init still finishes; the API still reports the instance running, Terraform still reports success, and your machine is up and unconfigured. A pipeline going green tells you the VM booted, not that it works.

Two habits that pay for themselves here:

  1. Make setup scripts re-runnable. A half-applied configuration is the normal failure, and the fix should be “run it again”, not “rebuild the box”.
  2. Fail loudly on purpose. Put set -euo pipefail at the top of a bootcmd/runcmd script and write a sentinel at the end — a file, a systemd unit reaching active, anything you can check from outside — so “did this work” has an answer that is not “SSH in and read the logs”.

Check what actually happened:

Terminal window
cloud-init status --long
sudo cloud-init analyze show
sudo journalctl -u cloud-final -b
sudo cat /var/log/cloud-init-output.log

If you cannot get in to run those, the VNC console is the way in — it works before the network does.