# Shelf Cloud — full API reference > Dedicated cores, unmetered bandwidth, honest pricing — priced live against the market, never above it. EC2-compatible API over hardware Shelf Cloud owns. Existing AWS SDKs and Terraform providers work unchanged against the endpoints below. Index: https://shelfcs.com/llms.txt # ACCOUNT — CLOUD ## Access keys Source: https://shelfcs.com/docs/account/cloud/access-keys How to get the EC2 access key pair that SigV4 signs with, where it really lives, and the one way it does not behave like AWS. ## Objective **Nothing on `ec2.shelfcs.com` works without an access key pair.** This page is how you get one, rotate one and withdraw one. `an internal service` (`listAccessKeys`, `createAccessKey`, `deleteAccessKey`), `an internal service`. ## What an access key actually is A **the identity service EC2 credential**. Not an IAM user, not an AWS-style key stored in an identity service of ours — a row the identity service holds against your cloud user, scoped to your project: ``` the identity service’s credential API ``` When you sign a request to `ec2.shelfcs.com`, the EC2 layer hands the signature to the identity service's the identity service’s signature-verification call to verify. the identity service is the authority on both your password and your access key. **Revoking the identity service user revokes both.** > [!primary] > > The [IAM access key actions](/docs/identity/iam/api-reference-access-keys) page > describes an AWS-shaped IAM service. That service is not deployed — there is > no `iam.shelfcs.com` in the > [published endpoint list](/docs/platform/endpoints/service-endpoints). What > exists today is what this page describes. Where the two disagree, this one is > what will answer your call. ## Getting one ### In the console Sign in and open **Access keys** (`/keys`). It is the same console, opened on that view. Mint, copy, withdraw — no CLI needed. ### Through the Account API | Method | Path | Returns | | --- | --- | --- | | `GET` | `/v1/access-keys` | Every key your cloud user holds | | `POST` | `/v1/access-keys` | `201` with a new `{ access, secret }` | | `DELETE` | `/v1/access-keys/{access}` | `204` | All three need a session (a Bearer JWT) **and an onboarded account**. Before onboarding they answer "Become a customer before making access keys." ```bash curl -X POST https://storefront.job-rss-processor.workers.dev/v1/access-keys \ -H "Authorization: Bearer $JWT" ``` ```json { "access": "…", "secret": "…" } ``` ## The one way it does not behave like AWS > [!warning] > > **`GET /v1/access-keys` returns the secret, not just the key id.** Amazon > shows a secret access key exactly once, at creation, and never again — the > guarantee being that AWS cannot show you something it does not keep in > recoverable form. > > the identity service does keep it, and this API returns it. Anyone who can obtain a > session token for your storefront account can read **every secret you hold**, > not merely mint a new one. Rotating your keys does not help if the session is > the thing that leaked. > > Treat your storefront login as equal in power to the keys themselves, and > treat a leaked JWT as a full compromise of the account's cloud credentials. The practical consequences: - A key you cannot find again is not lost. Read it back rather than minting a second one and leaving the first live. - "Show once" hygiene from AWS does not protect you here. Account hygiene does. ## Using one ```bash export AWS_ACCESS_KEY_ID=… export AWS_SECRET_ACCESS_KEY=… aws --endpoint-url https://ec2.shelfcs.com --region hel1 ec2 describe-instances ``` ```hcl provider "aws" { region = "hel1" access_key = var.ec2_access_key secret_key = var.ec2_secret_key skip_credentials_validation = true skip_requesting_account_id = true skip_metadata_api_check = true endpoints { ec2 = "https://ec2.shelfcs.com" } } ``` `region` must be `hel1`. It is part of the credential scope every signature carries, and a mismatch returns `SignatureDoesNotMatch` with nothing in the message that says why. ## Rotating There is no rotation endpoint, and no expiry. A key lives until you delete it. The safe order: 1. `POST /v1/access-keys` — mint the new pair. 2. Roll it out everywhere the old one is used. 3. Confirm the old one is idle. 4. `DELETE /v1/access-keys/{old access}`. Deleting is immediate and there is no grace period: the next signed request with that key fails. Delete after the rollout, never before. ## What a key can do Everything your cloud user can do, within your project — the `member` role, no more and no less. See [Quotas and limits](/docs/account/cloud/quotas-and-limits). There is **no per-key scoping**. A key is not narrower than the user that owns it: no read-only key, no key limited to one action, no key limited to one resource, no source-IP condition, no expiry. If you need a credential that can only list instances, this platform cannot yet give you one. Practically, that means a key handed to CI is a key that can terminate production. Until per-key scoping exists, the separation has to come from using different **accounts**, not different keys. ## How many No limit is set. Mint one per consumer — one for CI, one for your laptop, one for the tool that only reads — so that withdrawing one does not stop the others. The [`DescribeKeyPairs`](/docs/compute/ec2/ec2-api-key-pairs) SSH key pairs are a different thing entirely and are not affected by any of this. ## For internal teams Internal teams do not go through the storefront. A converging play mints one the identity service EC2 credential pair per tenant and writes both halves straight into Bitwarden, beside the cloud password — never printed to a console and never committed. A team receives an auth URL, a project name and a Bitwarden secret id. Rerunning the play is a converger: a tenant that already holds a credential gets nothing new, and a tenant missing one gets it back. ## Go further - [Account developer guide](/docs/account/guide/developer-guide) - [Account API](/docs/account/cloud/storefront-api-reference) - [Authentication](/docs/platform/conventions/authentication) - [Key pair actions](/docs/compute/ec2/ec2-api-key-pairs) — SSH keys, a different thing - [Service endpoints](/docs/platform/endpoints/service-endpoints) --- ## Account API Source: https://shelfcs.com/docs/account/cloud/storefront-api-reference Five endpoints — catalog, prices, me, onboard, billing portal. The front door between a signup and a cloud tenancy. ## Objective **This is the only API that turns a signup into a customer of the cloud.** Everything else on this platform assumes you already are one. It implements as little as possible: a hosted auth service owns identity, Stripe owns money, the identity service owns tenancy. The Account API is the glue, and the window that shows what is for sale. ## Endpoints ``` https://storefront.job-rss-processor.workers.dev ``` | Method | Path | Auth | Purpose | | --- | --- | --- | --- | | `GET` | `/v1/catalog` | **Public** | What we sell, price per hour, in stock | | `GET` | `/v1/prices` | **Public** | The price of every shape | | `GET` | `/v1/prices/comparison` | **Public** | What each shape costs elsewhere | | `GET` | `/v1/prices/rates` | **Public** | Every dimension of the bill at every other provider | | `GET` | `/v1/me` | Session | The caller's account | | `POST` | `/v1/onboard` | Session, verified email | Become a customer. Idempotent | | `GET` | `/v1/access-keys` | Session, onboarded | The caller's EC2 keys | | `POST` | `/v1/access-keys` | Session, onboarded | Mint one. `201` | | `DELETE` | `/v1/access-keys/{access}` | Session, onboarded | Withdraw one. `204` | | `POST` | `/v1/console/session` | Session, onboarded | A the identity service token for the web console | | `GET` | `/v1/usage` | Session | Machines running, and the month to date | | `POST` | `/v1/machines/{id}/console` | Session | A browser console URL for one machine | | `POST` | `/v1/billing/checkout` | Session | Stripe's page to put a card on file | | `POST` | `/v1/billing/card` | Session | Attach the card that page collected | | `GET` | `/v1/billing/portal` | Session, onboarded | Stripe's hosted portal URL, as JSON | > [!primary] > > The host above is the Worker's `workers.dev` address. **No domain is attached > yet** (`workers_dev = true`), so this hostname changes when one is. It is the > one endpoint on this platform that is not on `shelfcs.com`. > > `/v1/catalog`, `/v1/prices` and `/v1/prices/comparison` are cached for 60 > seconds at the edge; every other route is per-request. ## Authentication: a Bearer JWT from the hosted auth service Identity is **the hosted auth service**, the ledger database's managed Better Auth. This API and the site are on different hosts, so a browser never sends the hosted auth service's session cookie here — there is nothing to forward, and this service does no cookie handling at all. Instead: 1. The site asks the hosted auth service for a short-lived JWT (`POST /token` on Better Auth's own base path). 2. It sends it on every call as `Authorization: Bearer `. 3. This service verifies it against the hosted auth service's JWKS (EdDSA/Ed25519) and issuer. Default token expiry is **15 minutes**. Refresh from the hosted auth service; there is no refresh endpoint here. Three claims are read: `sub` (user id), `email`, `emailVerified`. A call to `/v1/onboard` with `emailVerified: false` is refused. ## `GET /v1/catalog` and `GET /v1/prices` Public — anyone can call them at any rate. Each answer still costs an admin the identity service token plus the compute service round trip (and, for prices, the rating service), so both are **cached for 60 seconds** at the edge, keyed on the request URL. A burst of shop traffic costs one round trip, not one per request. Use these to build a price page or a size picker without an account. See [How prices are set](/docs/billing/pricing/how-prices-are-set) for where the numbers come from. ## `POST /v1/onboard` Makes **five facts true, in order**, recording each before moving to the next, so a half-finished onboarding — a crash, a timeout, an upstream error — is finished by calling again: 1. **A Stripe customer**, tagged `metadata["tenant"] = cust-` — the name metering's tenant resolver depends on. 2. **A Stripe subscription** on the metered price, `collection_method: send_invoice`, 30 days to pay. An existing subscription is judged over every status but `canceled` and `incomplete_expired`, so a `past_due` subscription is still your one subscription. 3. **A the identity service project** named `cust-`, found by name before being created. 4. **A network of your own** — the networking service network, subnet (`10.200.0.0/24`, Quad9 DNS) and router with an external gateway, each looked up by name before being created. Without it `RunInstances` has nothing to attach a machine to. See [Networking](/docs/compute/ec2/ec2-api-networking). 5. **A the identity service user** in that project — your cloud login, from the request body — found before created, then granted the `member` role and a starting quota. Both are idempotent PUTs and are re-applied on every call until onboarding completes, so a retry that crashed between creating the user and granting its role still ends with both. **Finding before creating** is what makes a retry converge instead of hitting the identity service's 409 on a duplicate name. The request body carries your **cloud password**. It goes to the identity service — and it is also **kept by this service, in recoverable form**, so that `POST /v1/console/session` can sign you into the console without asking for it again. Read [Credentials and trust](/docs/platform/security/credentials-and-trust) before choosing it. > [!warning] > > The password is used **only the first time the user is created**. A later > `/v1/onboard` call with a different password does not change it, and there is > no password-reset endpoint yet. Choose it once, and keep it. ### Response The `Account` object, which once onboarded also carries: | Field | Meaning | | --- | --- | | `openstack_auth_url` | Your cloud identity endpoint — `https://identity.shelfcs.com` | | `openstack_username` | `cust-`, the name the identity service knows you by | Those two identify you to the cloud's own control APIs. Most customers never need them — the EC2 endpoint and your access keys are the supported path. See [Account developer guide](/docs/account/guide/developer-guide). ## `GET /v1/usage` Machines running and the month-to-date bill, in one call — the only cost-before-the-invoice surface this API has. It reads the compute service for what is running and the rating service for what it has cost so far, both scoped to your project. This is what the console's overview renders. There is still no cost explorer, no budget alert and no per-resource breakdown; for the rated rows behind the total, query the rating service directly — [Metering and invoicing](/docs/billing/pricing/metering-and-invoicing). ## Access keys `GET`, `POST` and `DELETE /v1/access-keys` mint and withdraw the EC2 credentials the AWS-shaped API signs with. They have a page of their own, including the one way they do not behave like AWS: [Access keys](/docs/account/cloud/access-keys). ## Consoles `POST /v1/console/session` returns `{ username, password, region, console_url }` — **your own cloud password, handed back to your own session**, so the page opens the console without asking you to type it. It requires an onboarded account that has a console password on file. What that implies: [Credentials and trust](/docs/platform/security/credentials-and-trust). `POST /v1/machines/{id}/console` returns a single-use, time-limited browser console URL for one machine, and refuses a machine that is not yours. See [Web console and VNC](/docs/platform/console/web-console-and-vnc). ## Cards `POST /v1/billing/checkout` returns a Stripe-hosted page for collecting a card; `POST /v1/billing/card` attaches the card that page collected, and rejects a request that does not name the checkout session it came from. Note that onboarding still creates a `send_invoice` subscription — a card on file does not by itself switch you to automatic collection. ## `GET /v1/me` The caller's account, including the same two the cloud platform fields once onboarded. Call it to find out whether onboarding has completed rather than assuming a `POST` succeeded. ## `GET /v1/billing/portal` Returns Stripe's hosted portal URL as JSON. Requires an onboarded account. The portal is where usage, invoices and payment method live — there is no billing UI of ours. See [Metering and invoicing](/docs/billing/pricing/metering-and-invoicing). ## Quotas set at onboarding **the compute service only** — cores and RAM. The `Quota` type carries no `volumes` or `gib` fields at all, because the block-storage service had no published hostname when this was written. It has one now (`volume.shelfcs.com`), so this is a gap that can close; until it does, a new project takes the block-storage service's deployment default rather than a quota chosen for it. the compute service's `the quota API` is itself the legacy quota API. the cloud platform's unified limits API replaces it eventually; that migration has not happened. ## What this API deliberately does not do - It does not handle a password for your *site* login — the hosted auth service does. It does hold your **cloud** password; see above. - It does not compute any amount of money — the rating service and Stripe do. - It knows nothing about internal teams. **There is one signup path, this one.** - There is no console (the web console) provisioning step: the console exists and is reached with the same the identity service credentials. > [!warning] > > The the identity service account this Worker authenticates as is the cloud **`admin`** > today, not a least-privilege service account. That account does not exist yet > (`the infrastructure repository`). It is a known gap in the > deployment, recorded here because it affects the blast radius of this service. ## Go further - [Access keys](/docs/account/cloud/access-keys) - [Account developer guide](/docs/account/guide/developer-guide) - [Quotas and limits](/docs/account/cloud/quotas-and-limits) - [How prices are set](/docs/billing/pricing/how-prices-are-set) - [Metering and invoicing](/docs/billing/pricing/metering-and-invoicing) - [Accounts and tenancy](/docs/platform/conventions/accounts-and-tenancy) --- ## Quotas and limits Source: https://shelfcs.com/docs/account/cloud/quotas-and-limits What a new account may consume, what sets it, what is not capped at all, and why the ceilings can never add up to more than the box. ## Objective **Read this before planning capacity.** A quota here is not a soft advisory number picked to look generous — the sum of every quota on the platform is asserted to fit inside what the hardware actually sells, and a change that would oversell the box fails before anything is touched. `an internal service` (`setQuota`), `an internal service`, `the infrastructure repository`. ## What a new account gets | Quota | Value | Set by | | --- | --- | --- | | vCPU (`cores`) | **4** | `POST /v1/onboard` | | Memory (`ram`) | **8192 MiB** (8 GiB) | `POST /v1/onboard` | | Volume count | the block-storage service deployment default | **Not set** | | Volume storage (GiB) | the block-storage service deployment default | **Not set** | Four vCPU and 8 GiB is two `cd-standard-2-4`, or one `cd-standard-4-16` minus the memory for it — enough to prove the platform works, not enough to run something in production without asking. Read your own, live: ``` cloud quota show cloud limits show --absolute ``` > [!warning] > > **Storage is not quota'd for customer accounts.** The onboarding path sets > the compute service's `cores` and `ram` only — the `Quota` type carries no `volumes` or `gib` > fields at all, because the block-storage service had no published hostname when it was written. > It has one now (`volume.shelfcs.com`), so this is a gap that can close. > > Until it does, a new project takes whatever the block-storage service's deployment default is, > which is a number nobody chose for it. Do not read "no quota" as "unlimited > storage": the pool is 1400 GiB total and shared with every instance root disk. ## Raising a quota There is no self-service quota API and no quota-increase form. Ask, and a person changes it. `PUT` on the compute service's `the quota API` is the call, and only an operator can make it. the compute service's `the quota API` is itself the **legacy** quota API — it is what the vendor SDK's `quotasets` package speaks and what the platform uses today. the cloud platform's unified limits API (registered limits plus project limits, shared across the compute service, the block-storage service and the networking service) replaces it eventually. That migration has not happened, and when it does, the call above changes. ## The ceiling that cannot be oversold Every internal tenant is declared with all four numbers: ```yaml - {name: platform-team, cores: 20, ram_mib: 40960, volumes: 10, gib: 900} - {name: data-team, cores: 2, ram_mib: 4096, volumes: 8, gib: 300} - {name: back-team, cores: 1, ram_mib: 3072, volumes: 4, gib: 100} - {name: front-team, cores: 1, ram_mib: 3072, volumes: 4, gib: 100} ``` The **first task** of the play that stamps them out is called "The ceilings must fit inside the offer". It sums every tenant's `cores`, `ram_mib` and `gib` and asserts the total fits inside what the catalog says the box sells. Adding a tenant, or raising an existing one, past that line fails the assert before a single resource is created. The fix is never to raise the ceiling. It is to shrink another allocation, or to free capacity. ## What the box sells | Resource | Capacity | Rule | | --- | --- | --- | | vCPU | 24 | 12 physical threads, CPU overcommitted at most 2× | | Memory | 51200 MiB (50 GiB) | **Never overcommitted.** 12 GiB stays with the host for the OS and the control plane | | Storage | 1400 GiB | Thin-provisioned, but the sum sold is capped so the pool cannot fill silently | | Guest addresses | 254 (`10.100.0.0/24`) | Well above what 24 vCPU can run | **Memory is the real limit.** CPU can be oversubscribed and is; memory cannot be, so 50 GiB is a hard ceiling on everything running at once. ## What has no limit at all Named so that their absence is deliberate, and so nobody plans around a control that does not exist: | | | | --- | --- | | API request rate | **No rate limiting.** None is configured at the ingress, and none on the authentication endpoint | | Volume IOPS and throughput | **Advertised, not enforced.** 48 IOPS and 11 MB/s are published per volume; no the block-storage service QoS applies them. One busy volume can take the pool | | Egress bandwidth | Unmetered, unbilled, unthrottled | | Snapshots | No count limit, no size limit, no expiry | | Instances per account | Bounded only by the vCPU and memory quota | | Key pairs, tags, security groups | Deployment defaults, not chosen | The first two are defects rather than features. Both are named on the [Volume types](/docs/storage/block/volume-types) and [Service endpoints](/docs/platform/endpoints/service-endpoints) pages, and neither should be relied on staying absent. ## When you hit one | Symptom | Cause | | --- | --- | | `InstanceLimitExceeded` on `RunInstances` | vCPU or memory quota | | `VolumeLimitExceeded` on `CreateVolume` | the block-storage service quota | | `InsufficientVolumeCapacity` | The **pool** is full, not your quota — nothing you can change | | `InsufficientInstanceCapacity` | The box has no room for that shape right now | The last two are capacity, not policy. A quota increase will not fix them, and we would rather return them than oversell and let two customers discover the same memory at once. ## Go further - [Account API](/docs/account/cloud/storefront-api-reference) — where the quota is set - [Instance types](/docs/compute/ec2/ec2-api-instance-types) — capacity and overcommit rules - [Volume types](/docs/storage/block/volume-types) — the storage pool - [Accounts and tenancy](/docs/platform/conventions/accounts-and-tenancy) --- ## Usage and cost Source: https://shelfcs.com/docs/account/cloud/usage-and-cost GET /v1/usage — what you are running and what the month has cost so far, the two hours of lag you must design around, and the machine console call beside it. ## Objective **This is the only way to see cost before an invoice exists.** There is no cost explorer, no budget alert, no per-resource breakdown and no tag allocation. ``` GET /v1/usage ``` Session required. Returns what your project is running, and what it has been rated so far this month. ## Response ```json { "machines": [ { "id": "…", "name": "web", "status": "ACTIVE", "flavor": "cd-standard-2-4", "addresses": ["10.200.0.7"], "created": "2026-09-01T09:14:02Z" } ], "month_to_date": 12.4, "since": "2026-09-01T00:00:00.000Z" } ``` | Field | Meaning | | --- | --- | | `machines[].id` | The machine's identifier in the control API | | `machines[].name` | | | `machines[].status` | `ACTIVE`, `SHUTOFF`, `BUILD`, `ERROR` — the control-API status, **not** the EC2 state name | | `machines[].flavor` | Our shape name | | `machines[].addresses` | Every address on the machine, private and public, flattened into one list | | `machines[].created` | | | `month_to_date` | Rated amount since `since`, in the billing currency | | `since` | **Midnight UTC on the first of the current month** | Neither number is computed by this service. What is running comes from the compute service; what it has cost comes from the rating service. This endpoint joins them and nothing more. ## The two-hour lag > [!warning] > > **An hour is rated about two hours after it ends.** The newest hours are not > in `month_to_date` yet. > > A machine you launched twenty minutes ago contributes **nothing** to this > figure. Neither does one you launched two hours ago, most likely. This is the single most important thing to know about the number. Consequences: - **Do not build an alert on a threshold** and expect it to fire promptly. By the time spend appears, it is at least two hours old. - **Do not use it to verify a shutdown worked.** Check `machines` for that — that half is live. - **Do not reconcile against it mid-hour.** Compare complete days. - **Month boundaries are UTC**, not your local midnight. A month-to-date figure read at 00:30 local time may still be counting the previous month depending on where you are. The `machines` list has no such lag. When you need "is it running", read that. When you need "what has it cost", accept that you are reading the recent past. ## Polling Rating happens hourly, so **poll hourly at most**. A tighter loop returns the same number and spends round trips to get it. If you want a live view of what is running, the `machines` half is safe to poll more often — but even then, prefer the compute API's own describe call, which is what it is for. ## `month_to_date` is a total, not a breakdown One number for the whole project. It does not tell you which machine, which shape, or which hour. For that, query the rating service directly with your cloud token — it returns the rated rows behind the total: which resource, which shape, which hour, which rate. That is the evidence to quote if a bill looks wrong. See [Metering and invoicing](/docs/billing/pricing/metering-and-invoicing). ## What is not in the number Only compute is rated. Storage, snapshots, egress and addresses contribute nothing to `month_to_date` because nothing rates them yet. **That is a gap, not a discount.** A cost model built on today's `month_to_date` will understate a future bill. See [How prices are set](/docs/billing/pricing/how-prices-are-set). ## Opening a machine's console ``` POST /v1/machines/{id}/console → { "url": "https://…" } ``` Returns a browser console for one machine — the equivalent of a Console button. Three properties worth knowing: 1. **Single-use and time-limited.** A fresh URL is minted on every call rather than stored, so do not cache one and do not paste one into a ticket. 2. **Ownership is checked first.** The machine is read and its project compared with yours before a console is opened; another project's machine returns a `404`. This check is the only thing separating customers here, because the service that performs it is privileged. 3. It works **before the network does** — which is the day you need it. See [Web console and VNC](/docs/platform/console/web-console-and-vnc). ## Example ```js const { machines, month_to_date, since } = await call("/v1/usage"); console.log(`${machines.length} running, €${month_to_date.toFixed(2)} since ${since}`); for (const m of machines) { console.log(m.name, m.flavor, m.status, m.addresses.join(", ")); } ``` Render `since` beside the figure. A cost with no period attached invites the reader to assume it means today. ## Go further - [Billing developer guide](/docs/billing/guide/developer-guide) - [Metering and invoicing](/docs/billing/pricing/metering-and-invoicing) - [Account API](/docs/account/cloud/storefront-api-reference) - [Web console and VNC](/docs/platform/console/web-console-and-vnc) --- # ACCOUNT — GUIDE ## Account developer guide Source: https://shelfcs.com/docs/account/guide/developer-guide Driving the account API from code — sessions and token refresh, idempotent onboarding, minting keys, reading usage, caching, CORS and error shapes. ## What this is How to call the account API from code. Concepts are in the [user guide](/docs/account/guide/user-guide); every endpoint and field is in the [API reference](/docs/account/cloud/storefront-api-reference). This API is **not** the compute API. It does not use SigV4 and it does not use access keys. ## Base URL ``` https://storefront.job-rss-processor.workers.dev ``` > [!primary] > > No custom domain is attached yet, so this hostname will change. Read it from > configuration rather than hard-coding it in a client you ship. ## Three levels of access | Level | What it needs | | --- | --- | | **Public** | Nothing. `/v1/catalog`, `/v1/prices`, `/v1/prices/comparison` | | **Session** | A Bearer JWT | | **Session + onboarded** | A Bearer JWT *and* a completed onboarding | A session call before onboarding does not fail with an auth error — it returns a `400` telling you to become a customer first. Handle that as a distinct state, not as a broken token. ## Sessions Identity is a hosted auth service, and this API and the site are on **different hosts** — so the browser's session cookie is never sent here and there is nothing to forward. Instead the site asks the auth service for a short-lived JWT and sends it on every call: ``` Authorization: Bearer ``` The token is verified against the issuer's JWKS on every request. Three claims are read: the user id, the email, and whether the email is verified. **Expiry is 15 minutes by default.** There is no refresh endpoint on this API — ask the auth service for a new token. Build the refresh into your client rather than treating a `401` as a dead end. ```js async function call(path, init = {}) { const jwt = await getFreshToken(); // yours; refresh when near expiry const res = await fetch(BASE + path, { ...init, headers: { ...init.headers, Authorization: `Bearer ${jwt}` }, }); if (!res.ok) throw await res.json(); return res.status === 204 ? null : res.json(); } ``` ## Onboarding, idempotently ```js await call("/v1/onboard", { method: "POST", headers: { "Content-Type": "application/json" }, body: JSON.stringify({ password: cloudPassword }), }); ``` Every step is find-before-create and every grant is re-applied until onboarding completes, so **calling it again is the recovery path** — not an error, and not something to guard against. If a call times out, call it again. Two things to get right: 1. **The body carries the cloud password.** It is used only on first creation; a later call with a different password does nothing. It is also **stored in recoverable form** so the console can be opened for the customer — see [Credentials and trust](/docs/platform/security/credentials-and-trust). 2. **Check completion with `GET /v1/me`** rather than assuming a `POST` succeeded. Once onboarded it also returns your cloud auth URL and username. ## Minting an access key ```js const key = await call("/v1/access-keys", { method: "POST" }); // 201 // { access, secret } await call("/v1/access-keys"); // list await call(`/v1/access-keys/${encodeURIComponent(access)}`, { method: "DELETE" }); // 204 ``` Rotate in this order: mint, roll out, confirm the old key is idle, delete. Deleting is immediate with no grace period. See [Access keys](/docs/account/cloud/access-keys) for the security property that differs from AWS. ## Reading usage ```js const usage = await call("/v1/usage"); ``` Machines running and the month-to-date cost, in one call. This is the only cost-before-the-invoice surface. **The cost half lags by about two hours** — an hour is rated roughly two hours after it ends — while the machines half is live. Field by field: [Usage and cost](/docs/account/cloud/usage-and-cost). ## Consoles ```js await call("/v1/console/session", { method: "POST" }); // web console await call(`/v1/machines/${id}/console`, { method: "POST" }); // one machine ``` The machine console returns a **single-use, time-limited URL** and refuses a machine that is not yours. Do not cache it, and do not paste one into a ticket. ## Caching The three public endpoints are cached for **60 seconds** at the edge, keyed on the URL. Every other route is per-request. That has one consequence worth coding around: `inStock` in the catalog can be up to a minute stale. It is a snapshot, not a reservation — another customer can take the capacity, and a launch can still fail with `InsufficientInstanceCapacity`. Treat the launch as the authoritative answer. ## Browsers and CORS Exactly **one origin** is allowed with credentials — this site's own. A page on another origin cannot call the session routes from a browser, by design. Server-side callers are unaffected. ## Errors Errors come back as JSON problem documents, not the EC2 XML shape the compute API uses. Do not share one error handler between the two. | Status | Means | Do | | --- | --- | --- | | `400` | Not onboarded, or a malformed body, or a card not matching its checkout session | Read the message; it names the missing precondition | | `401` | Missing, expired or unverifiable token | Refresh and retry once | | `403` | Email not verified | Verify, then retry | | `404` | Not yours, or no such thing | Treat as not-found | | `502`/`503` | An upstream did not answer — "the cloud did not answer about access keys", "prices are momentarily unavailable" | Retry with backoff. Nothing has been half-done: writes here are find-before-create | ## Go further - [User guide](/docs/account/guide/user-guide) - [API reference](/docs/account/cloud/storefront-api-reference) - [Access keys](/docs/account/cloud/access-keys) - [Catalog and availability](/docs/billing/pricing/catalog-and-availability) - [Quotas and limits](/docs/account/cloud/quotas-and-limits) --- ## Account user guide Source: https://shelfcs.com/docs/account/guide/user-guide From signing up to having a cloud tenancy — what onboarding creates, the two passwords, where the bill lives, and the one boundary you get. ## Two accounts, created together Signing up gives you a **site login**. Onboarding turns that into a **cloud tenancy**. They are separate systems and they do not share a session. | | Site login | Cloud login | | --- | --- | --- | | You are | Your email address | `cust-` | | Credential | Password, then a short-lived token | A separate cloud password, plus access keys | | Opens | This site, billing, the account API | Every cloud API, both consoles | You will end up holding **two passwords**. That is not a mistake in the design; it is one identity system for the shop and a different one for the cloud. ## What onboarding creates `POST /v1/onboard` — the button in the console — makes five things true, in order, recording each before moving to the next: 1. **A billing customer**, tagged with your tenant name. 2. **A subscription** on the metered price, invoiced with 30 days to pay. 3. **A project** — your tenancy, the boundary everything you own lives inside. 4. **A network of your own** — network, subnet (`10.200.0.0/24`) and router. Without it a launch would have nothing to attach a machine to. 5. **A cloud user** in that project, granted the `member` role and a starting quota. Because each step is recorded before the next begins, a half-finished onboarding — a crash, a timeout, an upstream error — is finished by pressing the button again rather than by starting over. > [!warning] > > **The cloud password is used only the first time.** Calling onboarding again > with a different password does not change it, and there is no password-reset > endpoint for it yet. Choose it once and store it somewhere you will still have > it in six months. > > It is also **kept in recoverable form** so the console can be opened for you > without asking. Use a password you use nowhere else, and read > [Credentials and trust](/docs/platform/security/credentials-and-trust). ## Verified email is required Onboarding refuses an unverified email address. Verify before you press it. ## Your first credentials Everything against the compute API needs an **access key pair**, which is not your password. Mint one in the console's Access keys view, or with one API call. Full detail — including that reading the list returns the secret, not just the key id — is in [Access keys](/docs/account/cloud/access-keys). ## What one account gets you | | | | --- | --- | | Projects | **One** | | Cloud users | **One** | | Roles | **One** — `member`, on your own project | | Networks | One, created for you | | Starting quota | 4 vCPU, 8 GB RAM | | Storage quota | Not set — see [Quotas and limits](/docs/account/cloud/quotas-and-limits) | **One account is one boundary.** You cannot make a second user for a colleague, cannot issue a read-only credential, and cannot separate staging from production inside it. If you need two environments that cannot touch each other, sign up twice. Raising a quota is a conversation, not a form: there is no self-service quota API. ## Where the bill lives There is no billing screen of ours. The account API hands you a link into the payment provider's hosted portal, and that is where usage, invoices and payment method live. Two things worth knowing: - **Only compute is billed.** Storage, snapshots, egress and addresses are not rated today. That is a gap, not a discount. - **`GET /v1/usage`** is the one place to see what is running and what the month has cost so far, before an invoice exists. See [Metering and invoicing](/docs/billing/pricing/metering-and-invoicing). ## Getting in when something is wrong - **Web console** — the whole cloud in a browser, signed in with your cloud credentials. - **Machine console** — a graphical console attached to one machine, which works before its network does. Neither has MFA; the cloud password is the only factor. See [Web console and VNC](/docs/platform/console/web-console-and-vnc). ## Go further - [Developer guide](/docs/account/guide/developer-guide) - [API reference](/docs/account/cloud/storefront-api-reference) - [Access keys](/docs/account/cloud/access-keys) - [Quotas and limits](/docs/account/cloud/quotas-and-limits) - [Compute user guide](/docs/compute/guide/user-guide) --- # BILLING — GUIDE ## Billing developer guide Source: https://shelfcs.com/docs/billing/guide/developer-guide Reading prices, stock and cost from code — which endpoint answers what, the null price that is not free, month-to-date usage, and the rated rows behind a total. ## What this is How to read prices and cost programmatically: building a price page, showing a customer what they will pay, or reconciling a bill. There is **no billing write API**. Nothing here changes what you owe. ## Which endpoint answers what | Question | Call | Auth | | --- | --- | --- | | What does every shape cost? | `GET /v1/prices` | **None** | | What is for sale, priced, and in stock right now? | `GET /v1/catalog` | **None** | | What does the same machine cost elsewhere? | `GET /v1/prices/comparison` | **None** | | What am I running, and what has this month cost? | `GET /v1/usage` | Session | | Where do I pay? | `GET /v1/billing/portal` | Session, onboarded | | Which rated rows make up that total? | The rating service directly | Cloud token | The first three are public and cached for 60 seconds at the edge. ## Building a price page ```bash curl -s $BASE/v1/catalog \ | jq '.[] | {name, vcpus, ramMiB, diskGiB, priceEurHour, inStock, aws: .aliases.aws}' ``` ```js const shapes = await (await fetch(`${BASE}/v1/catalog`)).json(); const sellable = shapes .filter(s => s.priceEurHour !== null) // see below .map(s => ({ ...s, monthly: s.priceEurHour * 730 })); ``` 730 hours is the month used in our own comparisons; use the same figure if you want your numbers to match ours. > [!warning] > > **`priceEurHour: null` does not mean free.** It means the rating system has no > mapping for that shape yet — it can be launched and nothing will bill for it. > That is a defect being fixed. Filter nulls out of a price page rather than > rendering "€0". ## `inStock` is a snapshot, not a reservation It is computed on every call from live capacity: in stock iff the remaining cores and memory fit one more of that shape. Two consequences for your code: 1. It can be up to **60 seconds stale** because of the cache, and another customer can take the capacity before you do. 2. **The launch is the authoritative answer.** Handle `InsufficientInstanceCapacity` from `RunInstances` even when the catalog said yes. Free capacity is clamped at zero, so you will never see negative stock even when the upstream accounting is transiently inconsistent. ## Comparison rows ```bash curl -s $BASE/v1/prices/comparison | jq '.comparisons[] | select(.shape=="cd-memory-2-16")' ``` Each row carries `provider`, `usd_hour`, `egress_included_gb`, `egress_usd_per_gb`, `usd_to_eur`, and — the part that matters — `source`, `as_of` and `method`. **Render the provenance.** A comparison without its source and date is a claim; with them it is checkable. The fields are there so you can show them. ## Month-to-date cost ```js const usage = await call("/v1/usage"); // session required ``` Machines running plus the month's cost so far, in one call. This is the only cost-before-invoice surface — there is no cost explorer, no budget alert, no per-resource breakdown and no tag allocation to build one from. Poll it hourly at most: metering runs on an hourly cron, so a tighter loop returns the same number. **An hour is rated about two hours after it ends**, so the newest hours are never in this figure. Full detail and what to design around: [Usage and cost](/docs/account/cloud/usage-and-cost). ## Rounding Amounts reach the payment provider in **minor units (cents)** as an integer. A rated amount under half a cent for a period rounds down to **zero** — the period is still reported so it is marked as seen, it simply bills nothing. If you are reconciling hour by hour, expect zero-valued periods and do not treat them as missing data. ## Timing Two delays stack. An hour is **rated about two hours after it ends**, and metering then emits at most a handful of periods per tick, so a backlog drains over several hours rather than at once. A cost figure is therefore **at least two hours behind reality**, and further behind when there is a backlog. Do not build an alert on a threshold that assumes real-time cost, and do not use cost to verify that a shutdown worked — use the running-machines list for that. ## Go further - [User guide](/docs/billing/guide/user-guide) - [Catalog and availability](/docs/billing/pricing/catalog-and-availability) - [How prices are set](/docs/billing/pricing/how-prices-are-set) - [Metering and invoicing](/docs/billing/pricing/metering-and-invoicing) - [Account developer guide](/docs/account/guide/developer-guide) --- ## Billing user guide Source: https://shelfcs.com/docs/billing/guide/user-guide What you pay for, what is currently free and why that will change, how the price is worked out, and where the invoice lives. ## What you are charged for **Compute, by the hour, by shape.** That is the whole of it today. | Resource | Billed | | --- | --- | | Instance, per hour | **Yes** | | Root disk | Included in the instance price | | Attached volumes | No | | Snapshots | No | | Egress bandwidth | No | | Addresses | No | | API requests | No | > [!warning] > > **Unbilled is not free.** Storage, snapshots and addresses are unrated because > the rating configuration does not cover them yet, not because we decided to > give them away. Closing those gaps will change bills. Do not build a cost > model that assumes storage stays at zero. Egress is different: it is deliberate. There is no egress line on the invoice and there is not meant to be one. ## How a price is worked out Nobody types a price in. For each shape: 1. Find the **cheapest current-generation instance at least as large** at the reference cloud — never smaller, so the comparison cannot flatter us. 2. Take its on-demand price in the reference region. 3. Subtract **20%**. 4. Convert to **euros** at the published rate. Burstable and ARM families are excluded from the comparison, because neither is equivalent: a burstable instance is not a fixed-core instance, and a different architecture is not the same machine. Including them would be picking the number that suits us. The full rule, the exclusions and what happens to a shape with no equivalent: [How prices are set](/docs/billing/pricing/how-prices-are-set). ## The price you see is the price that bills Prices come from the same rating system that produces your invoice, read live. There is no second price list beside it that could drift. You can read them without an account: ``` GET /v1/prices the price of every shape GET /v1/catalog prices, plus whether each shape is in stock right now GET /v1/prices/comparison what the same machine costs elsewhere ``` Every comparison row carries its source URL, the date it was read and the method used — so a comparison we publish can be checked rather than believed. ## When money moves | | | | --- | --- | | Metering | Hourly | | Aggregation | Summed over the period | | Collection | **Invoice, 30 days to pay** | | Currency | EUR | Card-on-file collection is a later phase; today you receive an invoice. You can add a card through the account settings, but that does not by itself switch you to automatic collection. **Small periods round down.** A rated amount under half a cent for an hour bills zero — the period is still recorded, it just charges nothing. ## Where to read it - **The hosted billing portal** — usage, invoices, payment method. There is no billing screen of ours; the account API hands you a link into the provider's. - **`GET /v1/usage`** — what is running and what the month has cost so far. This is the only way to see cost before an invoice exists. There is no cost explorer, no budget alert, no per-resource breakdown and no tag allocation. ## What every account must hold A subscription on the metered price. Without one there is nothing for usage to attach to — it is recorded but never invoiced. Onboarding creates it, and finds an existing one before creating, so a past-due subscription is still your one subscription rather than grounds for a second. ## Go further - [How a period is billed](/docs/billing/pricing/how-a-period-is-billed) — what an invoice line guarantees - [Developer guide](/docs/billing/guide/developer-guide) - [How prices are set](/docs/billing/pricing/how-prices-are-set) - [Metering and invoicing](/docs/billing/pricing/metering-and-invoicing) - [Catalog and availability](/docs/billing/pricing/catalog-and-availability) --- # BILLING — PRICING ## Catalog and availability Source: https://shelfcs.com/docs/billing/pricing/catalog-and-availability The two public endpoints behind the shop window — what is for sale, whether it is actually in stock right now, and what the same machine costs elsewhere. ## Objective **Both endpoints on this page are public.** No account, no key, no signature — build a price page, a size picker or a comparison against us without signing up. ``` GET https://storefront.job-rss-processor.workers.dev/v1/catalog GET https://storefront.job-rss-processor.workers.dev/v1/prices GET https://storefront.job-rss-processor.workers.dev/v1/prices/comparison ``` All three are cached for 60 seconds at the edge, keyed on the URL. A burst of traffic costs one round trip to the cloud, not one per request. ## `GET /v1/catalog` One entry per purchasable shape: | Field | Type | Meaning | | --- | --- | --- | | `name` | string | Our name, e.g. `cd-standard-2-8` | | `vcpus` | number | | | `ramMiB` | number | | | `diskGiB` | number | Root disk included with the shape | | `aliases` | object | Cloud → comma-joined names, e.g. `{"aws": "m5.large", …}` | | `inStock` | boolean | **Live.** See below | | `priceEurHour` | number or **null** | Euros per hour | **The catalog reads the cloud itself** — the compute service's flavors and the compute service's capacity — rather than a list maintained beside it. That is the point: it cannot disagree with what an order would actually get. A flavor that does not exist on the box is not in the catalog, and a shape the box cannot currently fit says so. ## `inStock` is computed, not declared ``` inStock ⟺ shape.vcpus ≤ freeVCPUs and shape.ramMiB ≤ freeRamMiB ``` In stock **iff the remaining capacity fits one more of it**. `freeVCPUs` and `freeRamMiB` come from the compute service's live hypervisor statistics on every call, not from a number someone updates. This is why the site can say it sometimes declines to quote and say why: a large shape goes out of stock before a small one, because the box genuinely cannot fit another. Two things to know when you consume it: 1. **It is a snapshot, not a reservation.** `inStock: true` means the capacity existed when the call was answered — up to 60 seconds ago, given the cache. Another customer can take it before you do. A launch can still fail with `InsufficientInstanceCapacity`; treat that as the authoritative answer. 2. **`vcpus_used` can exceed `vcpus`** on the real box — an inconsistent but observed upstream state, from overcommit accounting lag or a stale aggregate. Free capacity is clamped at zero rather than left negative, so nothing here ever reports negative stock, even transiently. ## `priceEurHour: null` means unpriced, not free A `null` price means **the rating service has no rating mapping for that flavor**. The shape exists and can be launched; nothing will bill for it. That is a defect in the rating configuration, not a discount, and it will be fixed. Do not build a cost model on a `null`. ## Where the price comes from Prices are read live from **the rating service's hashmap module**, in three hops: find the `instance` service, find its `flavor_name` field, read that field's `{flavor → cost}` mappings. This matters because it is the **same number that bills you**. Metering's ledger carries whatever the rating service rated, so the storefront reads prices from the rating service rather than keeping a copy — there is no second price list to drift out of step with the first. A cloud with no `instance` service configured in the hashmap module yields an empty map rather than an error: no price is a fact about the cloud, not a failure of the call. For the rule that decides what those numbers are, see [How prices are set](/docs/billing/pricing/how-prices-are-set). ## `GET /v1/prices/comparison` What the same machine costs elsewhere, one row per shape per provider: | Field | Meaning | | --- | --- | | `shape` | Our shape name | | `provider` | `aws`, `azure`, `ovh` | | `usd_hour` | Their price for the cheapest equivalent | | `egress_included_gb` | Their monthly allowance | | `egress_usd_per_gb` | Their rate beyond it | | `method` | How the figure was arrived at | | `source` | The URL it was read from | | `as_of` | The date it was read | | `usd_to_eur` | The rate used to convert | Every row carries `source`, `as_of` and `method`, so a comparison we publish can be checked rather than believed. The rows are written by `the infrastructure repository`' rating play, from the same computation that sets our own price — the cheapest current-generation AWS type at least as large — with the egress terms the catalog declares. **The storefront only reads them.** This is what the calculator on the pricing page runs on. ## Worked example ```bash curl -s https://storefront.job-rss-processor.workers.dev/v1/catalog \ | jq '.[] | select(.inStock) | {name, vcpus, ramMiB, priceEurHour, aws: .aliases.aws}' ``` ```bash curl -s https://storefront.job-rss-processor.workers.dev/v1/prices/comparison \ | jq '.comparisons[] | select(.shape=="cd-memory-2-16")' ``` ## What it will not tell you - **Not how much capacity is left**, only whether one more of a given shape fits. The absolute free vCPU and RAM figures are not published. - **Not your own quota** — that is `cloud quota show`, see [Quotas and limits](/docs/account/cloud/quotas-and-limits). - **Not storage price.** Attached volumes are not rated at all yet, so no volume appears in either endpoint. ## Go further - [How prices are set](/docs/billing/pricing/how-prices-are-set) - [Metering and invoicing](/docs/billing/pricing/metering-and-invoicing) - [Instance types](/docs/compute/ec2/ec2-api-instance-types) — the full alias table - [Account API](/docs/account/cloud/storefront-api-reference) --- ## How a period is billed Source: https://shelfcs.com/docs/billing/pricing/how-a-period-is-billed The guarantees behind your invoice — one row per hour, frozen once billed, never billed twice, and what happens when a rate changes after the fact. ## Objective **What you are actually promised about your bill.** Most clouds never write this down; the mechanics are worth knowing because they tell you what a disputed line can and cannot be. ## The unit is one project-hour The billing ledger holds **one row per project per hour** — the unit that is sold. A row records the hour it covers and the amount rated for it. Everything below is about the life of one such row. ## An unbilled hour converges; a billed hour is frozen | State | What a re-collection does | | --- | --- | | **Not yet billed** | Takes the rating system's latest word. A partially rated hour corrects itself | | **Already billed** | **Frozen.** The recorded amount never changes | That is the central rule: **what you were told is what the ledger keeps saying.** A number cannot be revised underneath an invoice you have already received. Rating an hour is not instantaneous, so an hour collected early may be rated incompletely. While it is unbilled it keeps converging on the true figure. Once it has been billed it stops moving, permanently. ## A zero-rated hour is not a sale An hour that rates to zero is **never recorded**. The ledger records what was sold, and nothing sold is not a row. If an hour was recorded and then walked back to zero **before** it was billed, the stale row is deleted outright — you are not invoiced for an hour that turned out to be nothing. If it had already been billed, the row is left exactly as it was. See the next section for what happens then. ## When a rate changes after you were billed This is the interesting case, and it is handled explicitly rather than swept up. If the rating system later reports a different amount for an hour you have already been billed for: 1. **The ledger row does not change.** You were told a number; that number stands. 2. **You are not silently re-billed**, and no correcting charge appears without anyone noticing. 3. **The discrepancy is recorded as an error** naming both the frozen amount you were billed and the new amount now reported, so a person reconciles it deliberately. The design choice worth noticing: a re-rate is neither applied silently nor discarded silently. Both of those would be easier. It is raised for a human, because a bill that changes on its own is worse than a bill someone has to look at. If you dispute a line, that record is what the dispute is resolved against. ## Never billed twice Before an hour is reported for billing it is marked as in flight. If a tick crashes between reporting and confirming, the next tick can see that the attempt was already made and does not report it a second time. The ledger is the guard, not the retry logic. That is why a failed run is safe to simply run again. ## When an hour fails to bill | Outcome | What happens | | --- | --- | | **Retryable failure** — the billing provider did not answer, or no customer record existed yet | The hour is marked as having failed and **rotated to the back of the queue**, so one stuck hour cannot block every other | | **Permanently unbillable** | The hour is **closed without being billed**, with the reason recorded | An hour becomes permanently unbillable when it falls outside the payment provider's own acceptance windows — an event too old to accept, or an unconfirmed report past the window in which its identifier can still be deduplicated. Those hours are **dropped, not billed** — and they are kept visibly distinct from hours that were billed normally, rather than being closed out to look the same. An hour that could not be charged is a fact about the system, and hiding it inside the "done" pile would make it unfindable. **In practice this means it is possible for usage to go uncharged.** It is recorded as such rather than quietly forgotten. ## Backlogs drain slowly, on purpose A tick reports only a handful of hours. A long backlog therefore takes several hours to clear rather than emptying at once. Combined with the rating delay, that is why [`GET /v1/usage`](/docs/account/cloud/usage-and-cost) can be more than two hours behind, and why a cost figure should never be treated as real-time. ## What this gives you - An invoice line, once issued, **means the same thing forever**. - You cannot be double-charged for an hour by a retry. - You cannot be charged for an hour that rated to nothing. - A change of mind after the fact is **surfaced, not applied**. - A failure to charge is **recorded, not hidden**. What it does not give you: real-time cost, a per-resource breakdown, or a guarantee that every hour is eventually charged. ## Go further - [Metering and invoicing](/docs/billing/pricing/metering-and-invoicing) — the chain end to end - [Usage and cost](/docs/account/cloud/usage-and-cost) — the lag, in detail - [How prices are set](/docs/billing/pricing/how-prices-are-set) — where the number comes from - [Billing user guide](/docs/billing/guide/user-guide) --- ## How prices are set Source: https://shelfcs.com/docs/billing/pricing/how-prices-are-set One rule, computed not typed — the cheapest current-generation AWS equivalent in London, less 20 percent, in euros. ## Objective **There is no price list maintained by hand.** Every instance price on this platform is derived from a published rule, recomputed on every run, so that the answer to "why does this cost that" is a calculation you can check rather than a number someone chose. `bootstrap/filter_plugins/pricing.py`, driven by `rating.yml`. ## The rule > The same as the closest competitor, less 20%. Precisely: 1. Take the shape (vCPU and RAM). 2. Find the **closest AWS current-generation instance type that is at least as big** — never smaller, so the comparison never flatters us. 3. Take its on-demand price in **`eu-west-2` (AWS London)**, the market we sell into. 4. Subtract the **discount: 20%**. 5. Convert to **EUR at the ECB reference rate**. `rating.yml` fetches the AWS spec feed and the FX rate once per run and hands both to `pricing.py`, which does the arithmetic. Nothing is cached across runs and nothing is typed in. | Field | Value | | --- | --- | | Rule | cheapest AWS equivalent, less the discount, in EUR | | Rated by | A running instance, by its shape, per hour | | Comparison region | `eu-west-2` | | Discount | 0.20 | | Currency | EUR | ## What the rule excludes, and why it matters `pricing.py` excludes the **burstable (`t`) and Graviton families** from the AWS comparison. Both would produce a cheaper "equivalent" that is not equivalent: a burstable instance is not a fixed-vCPU instance, and a Graviton instance is not x86. Comparing against them would be picking the number that suits us. That exclusion is also why four sub-2-GiB shapes were removed from the catalog: with `t` and Graviton excluded, AWS has no current-generation non-burstable type below 1 vCPU / 2 GiB, so every one of them priced identically to `cd-standard-1-2`. See [Instance types](/docs/compute/ec2/ec2-api-instance-types). ## How the inputs stay honest A computed price is only as trustworthy as the two numbers it is computed from — the competitor's price and the exchange rate. Both are handled so that a failure cannot quietly become a made-up price. **The exchange rate must answer on every run.** If the rate cannot be fetched, the pricing run **fails**. It does not fall back to the rate it used last time, and it does not leave the previously published prices in place while pretending they were recalculated. A price that says it was computed today was computed today, or it was not published. **The competitor spec feed is cached for 24 hours.** A stale cache still errs towards the last real fetch rather than towards an invented figure — the failure mode is a price up to a day old, never a price nobody measured. **Rules are written by converging, not by overwriting.** Each step reads the current rating rules before it writes, so re-running the process changes only what actually differs. Rerunning is safe and is how drift is corrected. **Nothing invents a number.** A shape with no competitor equivalent and no explicit price of its own is simply **not rated** — see below. The system would rather leave a gap than fill it with a plausible guess. ## Overrides An instance type may carry an explicit `price_eur_hour` of its own. **It wins** over the computed price when present. A shape priced neither way — no computed AWS equivalent and no explicit price — **is not rated and is not billed**. It is not free by intent; it is a hole, and a shape in that state should not be on sale. ## Reading the prices Prices are public and need no account: ``` GET https://storefront.job-rss-processor.workers.dev/v1/prices GET https://storefront.job-rss-processor.workers.dev/v1/catalog ``` `/v1/catalog` returns what we sell, the price per hour, and whether it is in stock. `/v1/prices` returns the price of every shape, and `/v1/prices/comparison` what the same machine costs elsewhere. All three are cached for 60 seconds. Field by field: [Catalog and availability](/docs/billing/pricing/catalog-and-availability). ## What is priced, and what is not | Resource | Rated today | | --- | --- | | Instance, per hour, by shape | **Yes** | | Root disk | Included in the instance price | | Attached volumes, per GiB-hour | **No** | | Snapshots | **No** | | Egress bandwidth | **No** | | Floating IP addresses | **No** | | API requests | **No** | Only compute is rated. The catalog's pricing rule prices the AWS-equivalent *instance type*, and there is no storage price rule — which is why the storage-dense shape `cd-storage-4-16-1000` was pulled from the catalog rather than sold at a compute price, and why a volume you create today costs nothing. **Unbilled is not free.** These are gaps in the rating configuration, not a pricing decision, and closing them will change bills. Do not build a cost model on the assumption that storage stays at zero. ## Egress We do not charge for egress today, and we do not yet promise not to. What the catalog does record is what the competition charges, with the date each figure was read: | Provider | Region | Included | Beyond that | Read on | | --- | --- | --- | --- | --- | | AWS | `eu-west-2` | 100 GB/month | $0.09 /GB | 2026-09-04 | | Azure | `uksouth` | 100 GB/month | $0.087 /GB | 2026-09-04 | | OVHcloud | FR public cloud | Effectively unlimited | $0 | 2026-09-04 | OVH is the honest exception: egress genuinely is included on their EU public cloud instances, the allowance is set high and the rate is zero, and saying otherwise on our own comparison would be a lie. Their whole price list is published without a key, in EUR per hour, so unlike AWS and Azure their base price is **read rather than modelled**. ## Currency and rounding Prices are quoted in **EUR**. Metering converts a rated amount to minor units (cents) before it reaches Stripe, and an amount under half a cent for a period **rounds down to zero** — the period is still recorded, it simply bills for nothing. See [Metering and invoicing](/docs/billing/pricing/metering-and-invoicing). ## Go further - [Metering and invoicing](/docs/billing/pricing/metering-and-invoicing) - [Instance types](/docs/compute/ec2/ec2-api-instance-types) - [Volume types](/docs/storage/block/volume-types) - [Metering events](/docs/platform/conventions/metering-events) - [Time, numbers and money](/docs/platform/conventions/time-numbers-and-money) --- ## Metering and invoicing Source: https://shelfcs.com/docs/billing/pricing/metering-and-invoicing the metering agent meters, the time-series store stores, the rating service rates, Stripe invoices — what happens to your usage every hour, and where you read the bill. ## Objective **Read this to know where your bill comes from and where to find it.** There is no billing screen of ours; there is a pipeline, and Stripe's hosted portal at the end of it. ## The chain ``` the metering agent → the time-series store → the rating service → metering Worker → the ledger database ledger → Stripe meters stores rates carries records invoices what runs the series per flavour the total it you ``` Every step is an the cloud platform component doing its own job. **Our own service computes nothing money-related** — it reads the rating service's rated total per project, writes it to a ledger, and tells Stripe. That is deliberate: the moment we start arithmetic on money, we own a class of bug we do not have to own. | Step | What it does | | --- | --- | | the metering agent | Meters what runs — instance existence, by flavour, over time | | the time-series store | Stores the time series | | the rating service | Rates the series against per-flavour prices from the catalog. Reachable at `rating.shelfcs.com` | | metering Worker | Cron, hourly. Reads the rating service's rated total per project | | the ledger database ledger | Records every period, whether or not it billed | | Stripe | Holds the subscription, sums the meter, sends the invoice | ## Cadence **Hourly.** A edge service on a cron trigger, not a daemon — nothing this team runs lives on its own machines. Each tick reads the rating service's `/v2/summary` for every project it meters, writes each period to the ledger, and emits the unemitted ones to Stripe. A period that fails to emit is retried on the next tick, because the ledger records it as unemitted rather than losing it. At most a handful of periods are emitted per tick, so a long backlog drains over several hours rather than in one.01 EUR per unit, monthly, metered | | Customer mapping | By `stripe_customer_id` in the event payload | | Collection | `send_invoice`, **30 days to pay** | The value is sent in cents because the Stripe meter payload takes an integer string. A rated amount under half a cent for the period **rounds down to 0**; that period is still reported — a zero value marks it seen — it simply bills for nothing. Card-on-file collection is a later phase. Today you receive an invoice. ## The tenant rule A the identity service project and its Stripe customer **share a name**, carried in the Stripe customer's `metadata["tenant"]`: ``` cust- ``` The metering service reads that mapping and never writes it. It is written once, at signup, by the storefront. This is the whole of the identity plumbing between your cloud resources and your bill — if that metadata is wrong, usage bills to the wrong customer, which is why nothing else is allowed to touch it. ## Where you read your bill **Stripe's hosted Customer Portal.** Ask the storefront for the URL: ``` GET /v1/billing/portal → { "url": "https://billing.stripe.com/..." } ``` It requires a session and an onboarded account, and the link it returns is short-lived. The portal shows usage and invoices, and it is where a payment method is managed. There is **no billing UI of ours**, and no billing API of ours beyond that one call. There is no cost-explorer, no budget alert, no per-resource cost breakdown, and no way to see today's accrued charge before the hour ticks. ## What every account must hold A subscription on the metered price. Without one there is nothing for a meter event to attach to and the usage is recorded but never invoiced. `POST /v1/onboard` creates it, and finds an existing one before creating — "existing" being judged over every subscription status except `canceled` and `incomplete_expired`, so a `past_due` invoice is still your one subscription and not grounds for a second. > [!warning] > > **Billing is in Stripe TEST mode.** The meter, product and price created on > 2026-09-03 are test-account objects, and the live ids replace them when > billing goes live. Nothing invoiced today is a real charge. Do not treat a > test invoice as evidence that live billing works. ## What is metered Compute, by flavour, per hour. Nothing else — see [How prices are set](/docs/billing/pricing/how-prices-are-set) for the full table of what is rated and what is not. Storage, snapshots, egress and addresses are all currently unbilled. ## Go further - [How prices are set](/docs/billing/pricing/how-prices-are-set) - [Metering events](/docs/platform/conventions/metering-events) - [Account API](/docs/account/cloud/storefront-api-reference) - [Accounts and tenancy](/docs/platform/conventions/accounts-and-tenancy) --- ## Provider rate feed Source: https://shelfcs.com/docs/billing/pricing/provider-rate-feed GET /v1/prices/rates — every dimension of the bill at every other provider, each rate carrying the vendor URL it was read from. ## Objective Most cloud comparisons compare one number: the hourly price of a machine. That is the number that flatters whoever wrote the comparison, because the rest of the bill is where the money is. This endpoint publishes **every dimension**, per provider, with the source of each figure. ``` GET https://storefront.job-rss-processor.workers.dev/v1/prices/rates ``` Public — no account, no key. Cached for 60 seconds. ## What it covers Per provider, a rate for each of: | Dimension | Why it is on the list | | --- | --- | | Machine | The number everyone quotes | | Block storage | Per GB-month, usually invisible until the invoice | | Egress | The line item that decides most bills | | Public IPv4 | Now charged by every major cloud | | Object storage | Storage plus request charges | | Managed database | The premium over running it yourself | ## Every rate carries its source Each rate carries **the vendor API URL it was read from**. Not a link to a pricing page a human transcribed — the endpoint the figure came out of. That is what makes the comparison checkable. You can fetch the same URL and see whether we read it correctly. ## Absent is not zero > [!primary] > > **What a vendor does not publish is absent, and the page names it rather than > assuming zero.** > > This matters more than it sounds. A comparison that treats an unpublished rate > as free silently makes a competitor look cheaper than they are, or lets us look > cheaper by omission. An absent dimension is shown as absent. So a `null` or a missing dimension means "this provider does not publish this", not "this is free". The one place a zero is real is where a provider genuinely includes something — egress on some EU public clouds is actually zero-rated with a high allowance, and saying otherwise would be a lie in our own favour. ## How the rows get there They are written by the rating play in the infrastructure repository, from the same computation that sets our own price. This service only reads them; it computes nothing. That is the same rule as everywhere else on this platform: one place decides a number, everything else reads it. ## Using it ```bash curl -s $BASE/v1/prices/rates | jq ``` Render the source and the date beside any figure you display. A comparison without its provenance is a claim; with it, it is evidence. ## Related endpoint `GET /v1/prices/comparison` answers a narrower question — what the **same machine** costs elsewhere, shape by shape. Use that for a per-shape comparison and this one for the full bill. See [Catalog and availability](/docs/billing/pricing/catalog-and-availability). ## Go further - [How prices are set](/docs/billing/pricing/how-prices-are-set) - [Catalog and availability](/docs/billing/pricing/catalog-and-availability) - [Billing developer guide](/docs/billing/guide/developer-guide) --- # Compute — EC2 ## Shelf Cloud EC2 API Reference Source: https://shelfcs.com/docs/compute/ec2/api-reference-welcome Endpoint, protocol, API version and how to read this reference ## Objective The Shelf Cloud EC2 API creates and manages virtual servers, their storage and their access keys. It implements the Amazon EC2 query API, so that clients written for Amazon EC2 work against it with only an endpoint change. **Read this reference to call the API directly, or to understand exactly what your SDK is sending.** ## Requirements - A Shelf Cloud account - An access key id and secret access key ## Instructions ### Endpoint ``` https://ec2.hel1.shelfcloud.com ``` One host per service per region. The service name and the region are part of every signature, so a client cannot be redirected to a differently-shaped host without every signature failing. **OPEN (Michael):** the public API domain. ### Protocol The EC2 query protocol. Parameters are sent as `application/x-www-form-urlencoded` in the body of a `POST`, or as a query string on a `GET`. Responses are XML. This is not a design choice we are free to revisit: it is what the AWS SDKs send, and matching it byte for byte is the entire reason existing tooling works. ### API version ``` Version=2016-11-15 ``` Every request carries the version. The SDKs supply it from their own service model; you do not set it by hand. Versions earlier than the implemented one are accepted, so that an older but still-supported SDK release continues to work. A version later than the implemented one is rejected with `InvalidParameterValue`. ### Authentication Every request is signed with AWS Signature Version 4. See the [authentication conventions](/docs/platform/conventions/authentication). The signing name for this service is `ec2`, and the signing region is the region in the endpoint host. ### Request ids Every response carries the request id in three places, because different clients look in different ones: - the `x-amzn-RequestId` response header, which every AWS SDK reads - a `` element in the body of a successful response - a `` element in the body of an error response The casing difference between the two body elements is inherited from Amazon EC2 and is preserved deliberately: a client that parses one and not the other must behave here exactly as it does there. Quote the request id in any support case. ### Reading an action page Each action page states: - what the action does, and what it changes - every request parameter, its type, and whether it is required - every response element - the errors the action can return, and what causes each - a worked example, as a signed request and its response Where an action is not idempotent, the page says so and says what a retry does. ## What is on the shelf | | | | --- | --- | | [Instance types](/docs/compute/ec2/ec2-api-instance-types) | Every shape we sell, and its AWS alias | | [Machine images](/docs/compute/ec2/ec2-api-machine-images) | The OS shelf, and the login user for each | | [Volume types](/docs/storage/block/volume-types) | What the disk underneath a volume actually is | | [Networking](/docs/compute/ec2/ec2-api-networking) | Default VPC, security groups, the shared guest network | | [How prices are set](/docs/billing/pricing/how-prices-are-set) | The rule every price is computed from | | [Service endpoints](/docs/platform/endpoints/service-endpoints) | Every hostname this cloud publishes | ## Go further - [Making requests](/docs/compute/ec2/api-reference-making-requests) - [Common errors](/docs/compute/ec2/api-reference-common-errors) - [Actions](/docs/compute/ec2/api-reference-actions) - [Account developer guide](/docs/account/guide/developer-guide) --- ## Common errors Source: https://shelfcs.com/docs/compute/ec2/api-reference-common-errors The error envelope, the HTTP status of each error class, and every shared error code ## Objective An error response is part of the contract. Callers branch on error codes, so a code whose meaning changes breaks working software as surely as a removed endpoint. **Read this guide to handle errors correctly, and to know which are worth retrying.** ## Requirements - None ## Instructions ### The error envelope Errors are returned as XML in the Amazon EC2 shape: ```xml InvalidInstanceID.NotFound The instance ID 'i-0a1b2c3d4e5f60718' does not exist b1e2c3d4-5678-90ab-cdef-1234567890ab ``` Note the element names: ``, ``, ``, and `` with that exact casing. There is no `` element; that belongs to the standard query protocol used by IAM, not to EC2. A client parsing EC2 errors expects this shape and no other. The `` is stable and is what you branch on. The `` is for a human and its wording may change between releases. Never parse the message. ### Status codes | Status | Meaning | | --- | --- | | `400` | The request was malformed, invalid, or expired | | `403` | The credentials are unknown, or policy denies the action | | `404` | Reserved; EC2 signals absence with a `.NotFound` code and a 400 | | `409` | The request conflicts with the state of the resource | | `500` | An error on our side | | `503` | Temporarily unavailable, or the request rate was exceeded | **We return no `401`.** Amazon EC2's own common-error list does document `NotAuthorized` at 401, and real EC2 returns 401 for some `AuthFailure` cases — so this is a deliberate divergence rather than parity, and it is stated as one. We return 403 for both. Two reasons: the AWS SDKs' automatic clock-skew correction is keyed on a 403 or 400 carrying a recognised code, and a 401 without a `WWW-Authenticate` header violates HTTP and confuses intermediaries that special-case it. **OPEN (Michael):** confirm the 403 choice against the SDK clock-skew handlers before shipping, since that behaviour is the stated reason for the whole table. ### Client errors | Code | Status | Cause | Retry | | --- | --- | --- | --- | | `AuthFailure` | 403 | The credentials could not be validated | No | | `IncompleteSignature` | 400 | The `Authorization` header is malformed | No | | `SignatureDoesNotMatch` | 403 | The signature does not match the request | No | | `InvalidClientTokenId` | 403 | The access key id does not exist or is inactive | No | | `MissingAuthenticationToken` | 403 | The request was not signed | No | | `RequestExpired` | 400 | The timestamp is outside the permitted skew | After fixing the clock | | `UnauthorizedOperation` | 403 | Policy does not allow this action | No | | `InvalidAction` | 400 | The action does not exist in this version | No | | `InvalidParameterValue` | 400 | A parameter has an unacceptable value | No | | `InvalidParameterCombination` | 400 | Two parameters cannot be used together | No | | `MissingParameter` | 400 | A required parameter is absent | No | | `InvalidPaginationToken` | 400 | The token is invalid or has expired | No | | `IdempotentParameterMismatch` | 400 | A client token was reused with different parameters | No | | `InvalidID` | 400 | An identifier is malformed | No | | `*.NotFound` | 400 | The named resource does not exist | Only if just created | | `*.Duplicate` | 400 | A resource with that name already exists | No | | `*.InUse` | 400 | The resource is in use and cannot be changed | No | | `DependencyViolation` | 400 | Another resource depends on this one | No | | `InstanceLimitExceeded` | 400 | An account quota would be exceeded | No | | `InsufficientInstanceCapacity` | 500 | No capacity for that type right now | Yes, with backoff | | `DryRunOperation` | 400 | `DryRun` was set; the call would have succeeded | No | | `Blocked` | 403 | The account is suspended | No | | `MalformedQueryString` | 404 | The query string is not well-formed | No | `InvalidParameterCombination` is returned, among other cases, when a describe call is given both a list of ids and `MaxResults`. That combination is rejected by Amazon EC2 and is rejected here. ### Server errors | Code | Status | Cause | Retry | | --- | --- | --- | --- | | `InternalError` | 500 | An unexpected failure on our side | Yes | | `Unavailable` | 503 | The service is temporarily unavailable | Yes | | `RequestLimitExceeded` | 503 | The request rate was exceeded | Yes, with backoff | Throttling is `RequestLimitExceeded` with a 503, which is what Amazon EC2 returns and what the Terraform AWS provider matches on for its backoff. A 429 would not be recognised by the installed base of older clients. ### Absence, denial, and what an error must not reveal A resource belonging to another account returns the same `.NotFound` as a resource that has never existed. Existence is not confirmed to a caller who is not entitled to know it. Within an account, a resource that exists but that policy forbids returns `UnauthorizedOperation`. The rule is therefore: - **another account, or no such resource** → `.NotFound` - **this account, denied by policy** → `UnauthorizedOperation` No error contains internal detail: no stack traces, no hostnames, no query text. ### Deprecating a code An error code is never given a new meaning. A behaviour needing a different meaning gets a new code, and the old one keeps returning what it always did until it is removed with notice. ## Go further - [Making requests](/docs/compute/ec2/api-reference-making-requests) - [Actions](/docs/compute/ec2/api-reference-actions) --- ## DescribeImages Source: https://shelfcs.com/docs/compute/ec2/ec2-api-describe-images List the images you can launch — parameters, filters, response elements, and what "the name rolls" means for a pipeline that pins an id. Lists the machine images available to you. Every image returned is one we publish. There are **no public images from other accounts, no marketplace and no customer-built images** — see [Machine images](/docs/compute/ec2/ec2-api-machine-images) for the shelf itself. Paginated. ## Request | Parameter | Type | Constraints | | --- | --- | --- | | `ImageId.N` | list | `ami-` followed by 8 or 17 hex characters | | `Filter.N` | list | See below | | `MaxResults` | integer | 5–1000 | | `NextToken` | string | Opaque | | `Owner.N` | list | Only `self` and our own account resolve. Others match nothing | | `ExecutableUsers.N` | list | Accepted; there are no shared images, so it never narrows anything | | `DryRun` | boolean | | Omitting `MaxResults` returns up to 1000 per page. ### Filters | Filter | Matches | | --- | --- | | `name` | The image name, e.g. `debian-13`. Wildcards allowed | | `image-id` | | | `architecture` | `x86_64` — the only architecture | | `state` | `available` | | `root-device-type`, `root-device-name` | | | `virtualization-type` | `hvm` | | `creation-date` | Wildcards allowed | | `description` | | | `tag:`, `tag-key` | | Filters naming things this platform never reports — product codes, billing products, image ownership by other accounts, deprecation times, TPM and boot-mode attributes — return `InvalidParameterValue` naming the filter, rather than matching nothing silently. ## Response | Element | Notes | | --- | --- | | `imageId` | `ami-…` | | `name` | `debian-13`, `ubuntu-24.04`, … | | `description` | **States the login user for that image** | | `imageState` | `available` | | `architecture` | `x86_64` | | `creationDate` | When we imported it | | `rootDeviceType`, `rootDeviceName` | | | `virtualizationType` | `hvm` | | `hypervisor` | | | `imageOwnerAlias` | Ours | | `public` | `true` — every image we publish is available to every account | | `blockDeviceMapping` | The root device and its minimum size | | `tagSet` | | > [!primary] > > **Read the login user out of `description`.** It is set by the distribution, > not by us, and using the wrong one produces a `Permission denied (publickey)` > that looks exactly like a broken key. That mistake costs more support time > than anything else on this platform. Elements the platform does not have — product codes, deprecation time, `imdsSupport`, `bootMode`, `tpmSupport`, `sriovNetSupport`, `enaSupport`, `platformDetails`, `usageOperation` — are omitted rather than invented. ## Names roll, ids change An image name points at the **newest build we have imported**. When we import a newer upstream build, the name moves to it and the `ami-` id changes. | If you | Then | | --- | --- | | Look up by name at deploy time | You get the newest build — usually what you want | | Pin a literal `ami-…` | You get that exact build until it is withdrawn | | Need bytes that cannot change for months | **Record the image checksum**, not the id | This is deliberate: upstream vendors publish rolling `latest` builds, and re-importing is how the shelf stays patched. > [!warning] > > **A stored `ami-` id is not a guarantee of identical bytes over time.** > Recording a fingerprint per image id is a known gap. Pin by checksum if > immutability matters to you. ## Errors | Code | Status | Cause | | --- | --- | --- | | `InvalidAMIID.Malformed` | 400 | The id is not the right shape | | `InvalidAMIID.NotFound` | 400 | No such image | | `InvalidParameterValue` | 400 | An unsupported filter, or `MaxResults` out of range | | `InvalidPaginationToken` | 400 | Expired or issued for different filters | ## Examples ```bash alias sc='aws --endpoint-url https://ec2.shelfcs.com --region hel1' sc ec2 describe-images \ --query 'Images[].[ImageId,Name,Description,CreationDate]' --output table sc ec2 describe-images --filters Name=name,Values='debian-*' ``` ```hcl data "aws_ami" "debian" { most_recent = true filter { name = "name" values = ["debian-13"] } } resource "aws_instance" "app" { ami = data.aws_ami.debian.id instance_type = "cd-standard-2-4" } ``` `most_recent = true` with a `name` filter is the right pattern here: the name is stable, the id behind it is not. ```python imgs = ec2.describe_images(Filters=[{"Name": "name", "Values": ["ubuntu-24.04"]}]) img = sorted(imgs["Images"], key=lambda i: i["CreationDate"])[-1] print(img["ImageId"], img["Description"]) # description names the login user ``` ## Not implemented `CreateImage`, `RegisterImage`, `CopyImage`, `DeregisterImage`, `ModifyImageAttribute` and `ResetImageAttribute` return `InvalidAction`. You cannot build or share an image on this platform yet. `DescribeImageAttribute` is available for the attributes the response above carries. ## Go further - [Machine images](/docs/compute/ec2/ec2-api-machine-images) — the shelf, with login users - [RunInstances](/docs/compute/ec2/ec2-api-run-instances) - [Instance types](/docs/compute/ec2/ec2-api-instance-types) - [Image and catalogue actions](/docs/compute/ec2/ec2-api-images-and-catalogue) --- ## DescribeInstances Source: https://shelfcs.com/docs/compute/ec2/ec2-api-describe-instances List and inspect instances — parameters, the filters that work and the ones that are refused, every response element, states, and pagination. Lists your instances, or inspects specific ones by id. Paginated. Read every page: a client that takes the first page and stops sees a partial answer with no error. ## Request | Parameter | Type | Constraints | | --- | --- | --- | | `InstanceId.N` | list | `i-` followed by 8 or 17 hex characters. Anything else: `InvalidInstanceID.Malformed` | | `Filter.N` | list | See below. Cannot be combined with `InstanceId.N` in a way that contradicts it | | `MaxResults` | integer | 5–1000. **Cannot** be combined with `InstanceId.N` | | `NextToken` | string | Opaque, from a previous response. Cannot be combined with `InstanceId.N` | | `DryRun` | boolean | | Omitting `MaxResults` gives up to **1000** instances per page, with a `nextToken` when there are more. ### Filters Filters AND together; the values within one filter OR together. Names and values are **case-sensitive**. Values may contain the wildcards `*` and `?`. Supported: | Filter | Matches | | --- | --- | | `instance-id` | | | `instance-type` | Our name | | `instance-state-name` | `pending`, `running`, `stopping`, `stopped`, `shutting-down`, `terminated` | | `instance-state-code` | | | `image-id` | | | `key-name` | | | `availability-zone` | | | `launch-time` | Wildcards allowed | | `private-ip-address` | | | `ip-address` | The public address | | `subnet-id`, `vpc-id` | | | `reservation-id` | | | `launch-index` | | | `client-token` | | | `root-device-name`, `root-device-type` | | | `architecture` | | | `security-group-id`, `security-group-name` | | | `block-device-mapping.*` | Attached volumes | | `network-interface.*` | The instance's interface | | `tag:`, `tag-key` | | > [!primary] > > **An unsupported filter is an error, not an empty result.** Filters naming > things this platform never reports — placement groups, IAM instance profiles, > metadata options, spot fields, dedicated hosts and the rest — return > `InvalidParameterValue` naming the filter. > > Matching nothing silently would be worse: your script would report zero > matches and carry on confidently with a wrong answer. ## Response Instances are grouped into reservations, exactly as on AWS — one reservation per `RunInstances` call. | Element | Notes | | --- | --- | | `reservationId` | `r-…` | | `instanceId` | `i-…` | | `imageId`, `instanceType`, `keyName` | `instanceType` is **our** name | | `instanceState` | `name` and `code` | | `privateIpAddress` | On your own subnet | | `ipAddress` | Public address, when one is associated | | `launchTime` | | | `placement.availabilityZone` | | | `amiLaunchIndex` | Position within its reservation | | `architecture`, `rootDeviceName`, `rootDeviceType` | | | `blockDeviceMapping` | Attached volumes and their attach state | | `groupSet` | Security groups | | `networkInterfaceSet` | The interface, its address and its subnet | | `tagSet` | | | `clientToken` | If one was sent at launch | | `stateReason`, `reason` | Why it reached its current state | | `monitoring` | Always `disabled` — there is no monitoring service | Elements this platform does not have are **omitted, never invented**. Every SDK treats an absent member as null, so nothing breaks; a fabricated value would. ## Instance states | State | Meaning | | --- | --- | | `pending` | Launching | | `running` | Booted. **Not necessarily configured** | | `stopping` / `stopped` | Off. Root disk and private address retained, storage still occupied | | `shutting-down` | Terminating | | `terminated` | Gone. Root disk destroyed | A terminated instance remains visible for a short window after termination, then disappears. That window is what lets a `terraform destroy` observe the terminal state it waits for instead of failing on a not-found. ## Ordering and pagination Reservations come back oldest first, by the creation time of their first instance; instances within a reservation by launch index. That is a stricter guarantee than AWS makes, which says order may vary — so code written against AWS is safe here. A reservation can span a page boundary and appear as a partial entry on each side. Join on `reservationId` if that matters to you. **An empty page can still carry a `nextToken`**, because filters are applied after a page is assembled. Stop when the token is absent, not when a page is empty. ## Errors | Code | Status | Cause | | --- | --- | --- | | `InvalidInstanceID.Malformed` | 400 | The id is not the right shape | | `InvalidInstanceID.NotFound` | 400 | No such instance — **or it is another account's**, which is byte-identical | | `InvalidParameterValue` | 400 | An unsupported filter, or `MaxResults` out of range | | `InvalidParameterCombination` | 400 | `MaxResults` or `NextToken` with `InstanceId.N` | | `InvalidPaginationToken` | 400 | Expired, invalid, or issued for different filters | A pagination token is tied to the filters and page size it was issued for. Changing either mid-iteration invalidates it. ## Examples ```bash alias sc='aws --endpoint-url https://ec2.shelfcs.com --region hel1' sc ec2 describe-instances \ --query 'Reservations[].Instances[].[InstanceId,InstanceType,State.Name,PrivateIpAddress]' \ --output table sc ec2 describe-instances --filters Name=tag:Name,Values=web Name=instance-state-name,Values=running ``` ```python for page in ec2.get_paginator("describe_instances").paginate( Filters=[{"Name": "instance-state-name", "Values": ["running"]}] ): for r in page["Reservations"]: for i in r["Instances"]: print(i["InstanceId"], i["InstanceType"]) ``` ```hcl data "aws_instances" "running" { instance_state_names = ["running"] } ``` ## Go further - [RunInstances](/docs/compute/ec2/ec2-api-run-instances) - [StartInstances, StopInstances, RebootInstances](/docs/compute/ec2/ec2-api-start-stop-instances) - [TerminateInstances](/docs/compute/ec2/ec2-api-terminate-instances) - [Pagination](/docs/compute/ec2/api-reference-pagination) - [Compute developer guide](/docs/compute/guide/developer-guide) --- ## EC2 API actions Source: https://shelfcs.com/docs/compute/ec2/api-reference-actions Every action the EC2 API implements, and every AWS action it does not ## Objective **Read this page to find out whether the call you want exists**, before writing code against it. An action listed here behaves as its Amazon EC2 counterpart behaves, unless this documentation states a difference. An action not listed here returns `InvalidAction`, and does so rather than silently doing something approximate. ## Requirements - Credentials with the relevant EC2 permissions ## Instructions ### Instances | Action | Purpose | Paginated | Idempotent | | --- | --- | --- | --- | | [`RunInstances`](/docs/compute/ec2/ec2-api-run-instances) | Launch instances | No | Yes, via `ClientToken` | | [`DescribeInstances`](/docs/compute/ec2/ec2-api-describe-instances) | List and inspect instances | Yes | Read | | [`TerminateInstances`](/docs/compute/ec2/ec2-api-terminate-instances) | Destroy instances | No | Naturally | | `StartInstances` | Start stopped instances | No | Naturally | | `StopInstances` | Stop running instances | No | Naturally | | `RebootInstances` | Restart in place | No | No | | `DescribeInstanceStatus` | Reachability and scheduled events | Yes | Read | | `ModifyInstanceAttribute` | Change a mutable attribute | No | Naturally | | `DescribeInstanceAttribute` | Read one attribute of an instance | No | Read | | [`DescribeInstanceTypes`](/docs/compute/ec2/ec2-api-instance-types) | List instance types and their shapes | Yes | Read | ### DryRun Every action above accepts a `DryRun` boolean. When set, permissions are evaluated and nothing is done: - the identity is permitted → `DryRunOperation` - the identity is not permitted → `UnauthorizedOperation` This is how a policy is tested before running a destructive command, and it is what the [policy actions guide](/docs/identity/iam/api-reference-policies) means by trying a call that should fail. ### Images | Action | Purpose | Paginated | | --- | --- | --- | | [`DescribeImages`](/docs/compute/ec2/ec2-api-machine-images) | List images | Yes | | `DescribeImageAttribute` | Read one attribute of an image | No | **OPEN (Michael):** whether `CreateImage`, `RegisterImage`, `CopyImage` and `DeregisterImage` exist in v1. They are the difference between customers using our images and customers building their own. ### Key pairs | Action | Purpose | Notes | | --- | --- | --- | | `CreateKeyPair` | Generate a pair; returns the private key once | Not idempotent | | `ImportKeyPair` | Register a public key you already hold | Not idempotent | | `DescribeKeyPairs` | List names and fingerprints | Never returns secrets | | `DeleteKeyPair` | Remove our copy of a public key | Naturally idempotent | ### Volumes | Action | Purpose | Paginated | | --- | --- | --- | | [`CreateVolume`](/docs/compute/ec2/ec2-api-volumes) | Create a volume | No | | [`DescribeVolumes`](/docs/compute/ec2/ec2-api-volumes) | List volumes | Yes | | `AttachVolume` | Attach a volume to an instance | No | | `DetachVolume` | Detach a volume | No | | `DeleteVolume` | Destroy a volume and its data | No | | `ModifyVolume` | Change a volume's size | No | | `DescribeVolumeStatus` | Volume health | Yes | ### Snapshots | Action | Purpose | Paginated | | --- | --- | --- | | [`CreateSnapshot`](/docs/compute/ec2/ec2-api-snapshots) | Point-in-time copy of a volume | No | | `DescribeSnapshots` | List snapshots | Yes | | `DeleteSnapshot` | Destroy a snapshot | No | ### Tags | Action | Purpose | | --- | --- | | `CreateTags` | Set tags on resources | | `DeleteTags` | Remove tags | | `DescribeTags` | List tags across resources | `CreateTags` sets tags to the values given. Repeating it with the same values changes nothing, which is why it needs no client token. ### Regions and zones | Action | Purpose | Paginated | | --- | --- | --- | | `DescribeRegions` | List regions | **No** | | `DescribeAvailabilityZones` | List zones in a region | **No** | | `DescribeAccountAttributes` | Account-level attributes | **No** | These three are deliberately unpaginated, because their Amazon EC2 models are unpaginated. Returning a token here would be discarded by every SDK, and the caller would silently receive a truncated list — the worst possible failure, because it looks like success. ### Networking EC2-Classic is disabled in the deployment (`disable_ec2_classic = True`), so every project gets a **default VPC** on first use and its security groups, instances and addresses are VPC ones. Full detail: [Networking](/docs/compute/ec2/ec2-api-networking). | Action | Purpose | | --- | --- | | `CreateSecurityGroup` | Create a security group | | `DescribeSecurityGroups` | List security groups | | `AuthorizeSecurityGroupIngress` | Add an ingress rule | | `AuthorizeSecurityGroupEgress` | Add an egress rule | | `RevokeSecurityGroupIngress` | Remove an ingress rule | | `RevokeSecurityGroupEgress` | Remove an egress rule | | `DeleteSecurityGroup` | Delete a security group | | `DescribeAddresses` | List addresses | **OPEN (Michael):** whether `AllocateAddress` and `AssociateAddress` are offered, which depends on how many public addresses the box has to give out. > [!primary] > > Each customer project gets its own network, subnet (`10.200.0.0/24`) and > router at onboarding, so another customer cannot reach your instances by > private address. Inside your own project everything is on one flat subnet; > security groups are the only tool for separating your own tiers. ### Not implemented in v1 Named so that their absence is deliberate: Spot instances, reserved instances, capacity reservations, placement groups, dedicated hosts, launch templates, fleets, Elastic Network Interfaces as first-class resources, customer-managed VPCs and subnets (a default VPC exists; managing your own does not), route tables, internet gateways, NAT gateways, VPN connections, transit gateways, VPC peering, network ACLs, IPv6, load balancers, instance metadata options, hibernation, instance recovery, EBS multi-attach, fast snapshot restore, `DescribeVolumesModifications`, `GetConsoleOutput`, `GetConsoleScreenshot`, the EBS direct APIs, and every GPU or accelerator attribute. There is also **no object storage service** of any kind on this platform — no Swift, no S3-compatible endpoint. Calling any of these returns `InvalidAction`. ## Go further - [RunInstances](/docs/compute/ec2/ec2-api-run-instances) - [DescribeInstances](/docs/compute/ec2/ec2-api-describe-instances) - [Volume actions](/docs/compute/ec2/ec2-api-volumes) - [Networking](/docs/compute/ec2/ec2-api-networking) - [Instance types](/docs/compute/ec2/ec2-api-instance-types) - [Machine images](/docs/compute/ec2/ec2-api-machine-images) - [Common errors](/docs/compute/ec2/api-reference-common-errors) --- ## Image and catalogue actions Source: https://shelfcs.com/docs/compute/ec2/ec2-api-images-and-catalogue DescribeImages, DescribeInstanceTypes, DescribeRegions, DescribeAvailabilityZones. ## Objective Four read actions that describe what is available to launch, and where. Three of them are **not paginated**, and that is a contract rather than an oversight — see below. ## Permissions All four require `"Resource": "*"`. Describe actions in EC2 are not resource-scoped. ## Instructions ### DescribeImages | Parameter | Type | Notes | | --- | --- | --- | | `ImageId.N` | list | Specific images | | `Filter.N` | list | `name`, `architecture`, `state`, `tag:` | | `MaxResults` | integer | 5 to 1000 | | `NextToken` | string | | **Paginated.** **Response:** for each image — `imageId`, `name`, `description`, `architecture`, `imageState`, `creationDate`, `rootDeviceType`, `blockDeviceMapping`. The `description` states the login user for that image. It is set by the distribution, not by us, and using the wrong one produces a permission denied that looks like a key problem. There are no public or shared images from other accounts. Every image returned is one we published. **Image names roll.** A name points at the newest build we have imported; when we import a newer one, the name moves and the id changes. Pin by `imageId` for a build that cannot change beneath you. ### DescribeInstanceTypes | Parameter | Type | Notes | | --- | --- | --- | | `InstanceType.N` | list | | | `Filter.N` | list | `vcpu-info.default-vcpus`, `memory-info.size-in-mib` | | `MaxResults` | integer | 5 to 100 | | `NextToken` | string | | **Paginated.** **Response:** `instanceType`, `vCpuInfo.defaultVCpus`, `memoryInfo.sizeInMiB`, `instanceStorageSupported`, `hypervisor`, `currentGeneration`. > [!primary] > > **Aliases do not appear in this response.** The Amazon EC2 model has no field > for them, and an SDK discards elements its model does not declare — so an > alias returned here would be invisible to the very command a customer runs to > look for it. > > The alias table is published in > [Instance types](/docs/compute/ec2/ec2-api-instance-types) instead. Each > alias is also accepted in `InstanceType.N` on this call, so a lookup by alias > works even though the response names our own type. ### DescribeRegions | Parameter | Type | Notes | | --- | --- | --- | | `RegionName.N` | list | | | `Filter.N` | list | `region-name`, `endpoint` | **Not paginated.** Never returns a `NextToken`. **Response:** `regionName`, `regionEndpoint`, `optInStatus`. ### DescribeAvailabilityZones | Parameter | Type | Notes | | --- | --- | --- | | `ZoneName.N` | list | | | `Filter.N` | list | `zone-name`, `state`, `region-name` | **Not paginated.** Never returns a `NextToken`. **Response:** `zoneName`, `zoneId`, `regionName`, `zoneState`. Each of our regions returns exactly one zone. Ask for it rather than constructing the name, and store what you are given: a caller that assumes one zone forever will need changing when there are two, and a caller that reads this call will not. ### Why three of these must not paginate `DescribeRegions`, `DescribeAvailabilityZones` and `DescribeAccountAttributes` are modelled as unpaginated in the Amazon EC2 service definition. Every AWS SDK therefore builds no paginator for them and **discards any `NextToken` we might return**. A caller would then receive a truncated list — half the zones, half the regions — with no error and no indication that anything was missing. Silent truncation of a location list is worse than an error, because the code proceeds confidently on a partial answer. So these actions return every result in one response, always. ## Go further - [Instance types](/docs/compute/ec2/ec2-api-instance-types) - [Machine images](/docs/compute/ec2/ec2-api-machine-images) - [Pagination](/docs/compute/ec2/api-reference-pagination) --- ## Instance metadata and user data Source: https://shelfcs.com/docs/compute/ec2/ec2-instance-metadata-and-user-data How a machine learns who it is — the config drive we force on every instance, what 169.254.169.254 does and does not do here, and why a green launch is not a working VM. ## Objective **Read this before writing cloud-init for this platform.** One deliberate difference from AWS changes how configuration reaches your machine, and it is not visible from the API side. ## Configuration arrives on a config drive, not over the network ``` force_config_drive = true ``` Every guest on this platform is handed its configuration on a **config drive** — a small read-only volume attached to the machine, labelled `config-2` — and never over the metadata network. the compute service normally attaches a config drive only when an instance asks for one, and the EC2 API has no parameter for asking. So no consumer of this cloud could ask; forcing it makes it the platform's behaviour rather than a per-launch accident. The reason is a dataplane one. A guest reaches `169.254.169.254` through the router of its own network — and a machine that has not been configured yet is exactly the machine least able to fix a fault on that path. The config drive is the only delivery that does not depend on the network coming up first. For Talos this matters twice over: it looks for the `config-2` volume first and falls back to the metadata address, so with this set it never touches the network to install itself. **What you should take from it:** if your machine did not get its keys or its user data, the network is not the first thing to check. Mount the drive and look. ```bash blkid -t LABEL=config-2 mkdir -p /mnt/cfg && mount /dev/disk/by-label/config-2 /mnt/cfg cat /mnt/cfg/openstack/latest/user_data cat /mnt/cfg/openstack/latest/meta_data.json ``` ## The metadata service There is one — `ec2-api-metadata` runs alongside the EC2 API — and it answers on the usual link-local address from inside an instance: ``` http://169.254.169.254/latest/meta-data/ http://169.254.169.254/latest/user-data ``` It is the EC2-shaped metadata service, so the familiar paths work: `instance-id`, `instance-type`, `local-ipv4`, `public-ipv4`, `hostname`, `placement/availability-zone`, `public-keys/0/openssh-key`. It is a fallback and a convenience here, not the delivery mechanism. Everything it serves is also on the config drive. ## What it does not do | AWS | Here | | --- | --- | | **IMDSv2 session tokens** (`PUT /latest/api/token`, `X-aws-ec2-metadata-token`) | **Not supported.** the compute service's metadata service has no session-token mode | | `HttpTokens=required` | Refused, not silently downgraded — `UnsupportedOperation` | | Instance tags in metadata (`InstanceMetadataTags`) | `disabled` | | IPv6 metadata endpoint | `disabled` | | IAM role credentials at `iam/security-credentials/` | **Nothing to serve.** There are no instance roles — see [Identity, as deployed](/docs/identity/keystone/identity-as-deployed) | The layer's defaults, which are the only values `MetadataOptions` will accept: ``` HttpTokens=optional HttpEndpoint=enabled HttpPutResponseHopLimit=1 HttpProtocolIpv6=disabled InstanceMetadataTags=disabled ``` `RunInstances` sends `MetadataOptions` only when Terraform has a `metadata_options` block; any member that differs from the above is rejected rather than quietly ignored. > [!warning] > > **The metadata endpoint is unauthenticated from inside the guest, and > `HttpTokens=required` cannot be set.** IMDSv2 exists on AWS specifically to > blunt SSRF — an application tricked into fetching a URL cannot reach IMDS > because it will not send the token header. That defence is not available here. > > There are no role credentials to steal, which is the worst of what SSRF gets > on AWS. But your SSH public key, hostname and network layout are readable by > anything that can make an HTTP request from inside the machine. ## User data | | | | --- | --- | | Parameter | `UserData` on `RunInstances` | | Encoding | Base64 | | Limit | **16384 bytes decoded** — AWS's limit, enforced, even though the compute service would allow 65535 | | Logging | Never logged. The model marks the shape sensitive | | Delivery | Config drive, and the metadata service | ```bash aws --endpoint-url https://ec2.shelfcs.com --region hel1 \ ec2 run-instances --image-id ami-… --instance-type cd-standard-2-4 \ --key-name mykey --user-data file://cloud-init.yaml ``` ```hcl resource "aws_instance" "app" { ami = data.aws_ami.debian.id instance_type = "cd-standard-2-4" key_name = aws_key_pair.mine.key_name user_data = file("cloud-init.yaml") } ``` Over 16384 bytes decoded is `InvalidParameterValue`. If your configuration is bigger than that, put a fetch in the user data and the payload somewhere your machine can reach — noting that there is no object storage on this platform to put it in. ## The guest environment cloud-init lands in | | | | --- | --- | | Login user | Set by the image, never by us — `debian`, `ubuntu`, `rocky`, `almalinux`, `fedora`, `opensuse`, `arch`. See [Machine images](/docs/compute/ec2/ec2-api-machine-images) | | Root SSH | Disabled by the vendor image. Your key goes to the login user | | Address | On `10.100.0.0/24`, gateway and NAT at `10.100.0.1` | | DNS | Quad9 — `9.9.9.9`, `149.112.112.112` | | Root disk | 20–160 GB by shape, deleted with the instance | Talos takes no cloud-init: it reads Talos machine configuration from the same user-data channel, and has no SSH or login user at all. ## A green launch is not a working machine > [!warning] > > **cloud-init's `runcmd` swallows failures.** Every command in a `runcmd` block > can fail and cloud-init still finishes; the API still reports the instance > `running`, Terraform still reports success, and your machine is up and > unconfigured. A pipeline going green tells you the VM booted, not that it > works. Two habits that pay for themselves here: 1. **Make setup scripts re-runnable.** A half-applied configuration is the normal failure, and the fix should be "run it again", not "rebuild the box". 2. **Fail loudly on purpose.** Put `set -euo pipefail` at the top of a `bootcmd`/`runcmd` script and write a sentinel at the end — a file, a systemd unit reaching `active`, anything you can check from outside — so "did this work" has an answer that is not "SSH in and read the logs". Check what actually happened: ```bash cloud-init status --long sudo cloud-init analyze show sudo journalctl -u cloud-final -b sudo cat /var/log/cloud-init-output.log ``` If you cannot get in to run those, the [VNC console](/docs/platform/console/web-console-and-vnc) is the way in — it works before the network does. ## Go further - [RunInstances](/docs/compute/ec2/ec2-api-run-instances) — the `UserData` and `MetadataOptions` parameters in full - [Machine images](/docs/compute/ec2/ec2-api-machine-images) — login users - [Key pair actions](/docs/compute/ec2/ec2-api-key-pairs) - [Networking](/docs/compute/ec2/ec2-api-networking) — the guest network - [Web console and VNC](/docs/platform/console/web-console-and-vnc) --- ## Instance types Source: https://shelfcs.com/docs/compute/ec2/ec2-api-instance-types Every shape we sell, its size, its AWS alias, and the equivalents on GCP, Azure, Alibaba, IBM and OVH. ## Objective **The alias table lives here** — `DescribeInstanceTypes` cannot carry it, because the Amazon EC2 model has no field for an alias and an SDK discards elements its model does not declare. is the only source of truth for the offer; `bootstrap/roles/openstack/tasks/resources.yml` creates one the compute service flavor per row and carries the aliases as flavor metadata (`aws:alias`, `gcp:alias`, `azure:alias`, `alibaba:alias`, `ibm:alias`). ## Naming ``` cd--- ``` `cd` is compute desk. `cd-standard-2-8` is 2 vCPU and 8 GiB. Sub-GiB shapes would carry an `m` suffix and be MiB; none are sold today. **An alias is never a lie about size.** An AWS instance-type name appears in the `aliases` column only where its vCPU and RAM match ours exactly, so existing Terraform that says `m5.large` gets 2 vCPU and 8 GiB, which is what `m5.large` means. Where no current-generation, non-burstable AWS name matches a shape, the shape carries no AWS alias at all rather than a near-enough one. ## The shapes ### Small — one vCPU Every modern AWS family starts at 2 vCPU, so there is no honest current-generation, non-burstable AWS name for a 1-vCPU shape. None of these carry an AWS alias. | Name | vCPU | RAM | Root disk | Azure | OVH | | --- | --- | --- | --- | --- | --- | | `cd-standard-1-2` | 1 | 2 GiB | 20 GiB | `Standard_B1ms` | `d2-2` | | `cd-standard-1-4` | 1 | 4 GiB | 40 GiB | | `d2-4` | | `cd-memory-1-8` | 1 | 8 GiB | 40 GiB | | | ### General purpose — 1:2 and 1:4 | Name | vCPU | RAM | Root disk | AWS | GCP | Azure | Alibaba | IBM | OVH | | --- | --- | --- | --- | --- | --- | --- | --- | --- | --- | | `cd-standard-2-2` | 2 | 2 GiB | 40 GiB | | `e2-highcpu-2` | | | | | | `cd-standard-2-4` | 2 | 4 GiB | 80 GiB | | | | | | `c3-4` | | `cd-standard-2-8` | 2 | 8 GiB | 80 GiB | `m5.large` | `e2-standard-2`, `n2-standard-2` | `Standard_D2s_v5`, `Standard_B2ms` | `ecs.g6.large` | `bx2-2x8` | `d2-8` | | `cd-standard-4-16` | 4 | 16 GiB | 120 GiB | `m5.xlarge` | `e2-standard-4`, `n2-standard-4` | `Standard_D4s_v5` | `ecs.g6.xlarge` | `bx2-4x16` | `b3-16` | | `cd-standard-8-32` | 8 | 32 GiB | 160 GiB | `m5.2xlarge` | `e2-standard-8`, `n2-standard-8` | `Standard_D8s_v5` | `ecs.g6.2xlarge` | `bx2-8x32` | `b3-32` | `cd-standard-2-2` is the **default instance type**: it is what a launch that names no type receives. ### Compute — 1:2 | Name | vCPU | RAM | Root disk | AWS | GCP | Azure | Alibaba | IBM | OVH | | --- | --- | --- | --- | --- | --- | --- | --- | --- | --- | | `cd-compute-2-4` | 2 | 4 GiB | 80 GiB | `c5.large` | `c2d-highcpu-2` | `Standard_F2s_v2`, `Standard_B2s` | `ecs.c6.large` | `cx2-2x4` | `c3-4` | | `cd-compute-4-8` | 4 | 8 GiB | 120 GiB | `c5.xlarge` | `c2d-highcpu-4` | `Standard_F4s_v2` | `ecs.c6.xlarge` | `cx2-4x8` | `c3-8` | | `cd-compute-8-16` | 8 | 16 GiB | 160 GiB | `c5.2xlarge` | `c2d-highcpu-8` | `Standard_F8s_v2` | `ecs.c6.2xlarge` | `cx2-8x16` | `c3-16` | ### Memory — 1:8 | Name | vCPU | RAM | Root disk | AWS | GCP | Azure | Alibaba | IBM | OVH | | --- | --- | --- | --- | --- | --- | --- | --- | --- | --- | | `cd-memory-2-16` | 2 | 16 GiB | 80 GiB | `r5.large` | `e2-highmem-2` | `Standard_E2s_v5` | `ecs.r6.large` | `mx2-2x16` | `r3-16` | | `cd-memory-4-32` | 4 | 32 GiB | 120 GiB | `r5.xlarge` | `e2-highmem-4` | `Standard_E4s_v5` | `ecs.r6.xlarge` | `mx2-4x32` | `r3-32` | **Verification.** AWS aliases are checked against live AWS spec data on every deploy, and an alias that does not match fails it — see [What we verify before selling it](/docs/platform/security/verified-claims). GCP, Azure, Alibaba and IBM equivalents are exact-size matches pinned by hand, because no public feed is wired for them yet — they are a convenience for comparison, not a contract. ## Using an alias An alias is accepted anywhere an instance type is: ``` aws --endpoint-url https://ec2.shelfcs.com --region hel1 \ ec2 run-instances --instance-type m5.large --image-id ami-… --count 1 ``` ```hcl resource "aws_instance" "app" { instance_type = "m5.large" # resolves to cd-standard-2-8 ami = data.aws_ami.debian.id } ``` The **response names our type, not the alias**: `DescribeInstances` returns `cd-standard-2-8` for an instance launched as `m5.large`. Terraform records that in state, so the next plan shows a diff on `instance_type` unless you either write our name in the configuration or add `lifecycle { ignore_changes = [instance_type] }`. Writing our name is the better answer. `DescribeInstanceTypes` accepts an alias in `InstanceType.N` — a lookup by alias works — and returns our name in `instanceType`. ## What no shape has Named so their absence is deliberate: - **No burst credits.** There is no `t`-family equivalent and no CPU credit balance. A shape's vCPU count is what it always gets. - **No GPU, no accelerator of any kind.** - **No local NVMe or instance store.** `instanceStorageSupported` is `false` on every type. Storage is the root disk plus attached volumes — see [Volume types](/docs/storage/block/volume-types). - **No bare metal, no dedicated host, no placement group.** ## Capacity and overcommit | Rule | Value | | --- | --- | | Physical | 12 threads, 62 GB RAM, 1.7 TB HDD | | vCPU capacity sold | 24 | | CPU overcommit | at most 2× | | Memory capacity sold | 51200 MiB (50 GiB) | | Memory overcommit | **never** | | Host reserve | 12 GiB — OS plus the control plane, headroom included | **RAM is never overcommitted.** If a shape says 8 GiB, 8 GiB of real memory is set aside for it. CPU is overcommitted at most 2×, which is why a busy neighbour can cost you cycles but never memory. The box is IO-bound before it is CPU-bound. Read [Volume types](/docs/storage/block/volume-types) before sizing anything that touches disk. ## Shapes that were removed, and why Four shapes below `cd-standard-1-2` — `cd-nano-1-128m`, `cd-nano-1-256m`, `cd-nano-1-512m`, `cd-standard-1-1` — are gone. AWS has no non-burstable, current-generation type under 1 vCPU / 2 GiB, and the pricing rule excludes the burstable (`t`) and Graviton families that do go that small. Every one of those four therefore resolved to the *same* AWS-equivalent price as `cd-standard-1-2`: four SKUs charging the same money for strictly less machine. `cd-storage-4-16-1000` — the storage-dense shape an HDD box is actually good at — is also removed, for a different reason: the pricing rule prices the AWS-equivalent instance type, not attached storage, so there is no honest way to price it yet. It comes back when a storage price rule exists. ## Go further - [Machine images](/docs/compute/ec2/ec2-api-machine-images) - [Image and catalogue actions](/docs/compute/ec2/ec2-api-images-and-catalogue) — `DescribeInstanceTypes` itself - [How prices are set](/docs/billing/pricing/how-prices-are-set) - [RunInstances](/docs/compute/ec2/ec2-api-run-instances) --- ## Key pair actions Source: https://shelfcs.com/docs/compute/ec2/ec2-api-key-pairs CreateKeyPair, ImportKeyPair, DescribeKeyPairs, DeleteKeyPair. ## Objective Key pairs are how SSH access to an instance is established. We hold public keys only. ## Permissions | Action | Resource scope | | --- | --- | | `ec2:CreateKeyPair` | `key-pair/*` | | `ec2:ImportKeyPair` | `key-pair/*` | | `ec2:DescribeKeyPairs` | Requires `"Resource": "*"` | | `ec2:DeleteKeyPair` | `key-pair/` | ## Instructions ### CreateKeyPair Generates a pair and returns the private key **once**. | Parameter | Type | Required | Notes | | --- | --- | --- | --- | | `KeyName` | string | Yes | Unique within the account, 1–255 characters | | `KeyType` | string | No | `ed25519` or `rsa`. **PROPOSED:** default `ed25519` | | `TagSpecification.N` | list | No | | | `DryRun` | boolean | No | | **Response:** `keyName`, `keyFingerprint`, `keyPairId`, and `keyMaterial` — the PEM-encoded private key. ```xml b1e2c3d4-5678-90ab-cdef-1234567890ab laptop key-0a1b2c3d4e5f60718 1f:51:ae:28:bf:89:e9:d8:1f:25:5d:37:2d:7d:b8:ca -----BEGIN PRIVATE KEY----- ... -----END PRIVATE KEY----- ``` > [!warning] > > `keyMaterial` appears in this response and nowhere else. We do not store it in > a recoverable form. Losing it means every instance launched with this key pair > becomes unreachable, and there is no console through which to recover one. **Not idempotent.** No `ClientToken` exists on this action, so a retry after a lost response generates a *second, different* key pair — and the private key from the first response is gone. Prefer `ImportKeyPair`, which is safe to retry. **Errors** | Code | Status | Cause | | --- | --- | --- | | `InvalidKeyPair.Duplicate` | 400 | That name already exists | | `InvalidParameterValue` | 400 | Name too long, or an unsupported key type | | `KeyPairLimitExceeded` | 400 | A quota would be exceeded | ### ImportKeyPair Registers a public key you already hold. The recommended path: your private key never crosses the network. | Parameter | Type | Required | Notes | | --- | --- | --- | --- | | `KeyName` | string | Yes | | | `PublicKeyMaterial` | blob | Yes | Base64 of an OpenSSH or RFC 4716 public key | | `TagSpecification.N` | list | No | | **Response:** `keyName`, `keyFingerprint`, `keyPairId`. No key material, because we never had the private half. Importing the same key under the same name twice returns `InvalidKeyPair.Duplicate`; importing the same key under a different name is permitted, and produces two key pairs with the same fingerprint. ### DescribeKeyPairs | Parameter | Type | Notes | | --- | --- | --- | | `KeyName.N` | list | | | `KeyPairId.N` | list | | | `Filter.N` | list | `key-name`, `fingerprint`, `tag:` | **Response:** for each, `keyName`, `keyPairId`, `keyFingerprint`, `createTime`, `tagSet`. **There is no action that returns a private key.** None exists to be called, in this API or any other, because we hold none. Not paginated — key pair counts are bounded by a quota small enough that paging would add complexity without protecting anything. ### DeleteKeyPair | Parameter | Type | Required | | --- | --- | --- | | `KeyName` or `KeyPairId` | string | Yes, one of them | Removes our copy of the public key. **It does not remove the key from running instances.** The public key was written into the instance at first boot and lives in that instance's `authorized_keys`. Anyone holding the private key keeps access after deletion. To actually revoke access to a running instance, remove the key from `authorized_keys` inside the instance. Deleting the key pair only prevents future launches from using it. Deleting a key pair that does not exist **succeeds** — this action is one of the few that does not return not-found, matching Amazon EC2's behaviour. ## Go further - Key pairs - [RunInstances](/docs/compute/ec2/ec2-api-run-instances) --- ## Machine images Source: https://shelfcs.com/docs/compute/ec2/ec2-api-machine-images The eleven images on the shelf, the login user for each, the minimum root disk, and what "the name rolls" means for your pipeline. ## Objective **Read this page to pick an image and to find its login user.** Using the wrong user produces a permission-denied that looks exactly like a broken key, and costs an hour. **official vendor cloud image**, downloaded from the vendor's own URL and imported into the image service. We build none of them, and we modify none of them. ## The shelf | Name | Login user | Minimum root disk | Format | | --- | --- | --- | --- | | `debian-12` | `debian` | 4 GiB | qcow2 | | `debian-13` | `debian` | 4 GiB | qcow2 | | `ubuntu-22.04` | `ubuntu` | 8 GiB | qcow2 | | `ubuntu-24.04` | `ubuntu` | 8 GiB | qcow2 | | `rocky-9` | `rocky` | 10 GiB | qcow2 | | `alma-9` | `almalinux` | 10 GiB | qcow2 | | `fedora-42` | `fedora` | 10 GiB | qcow2 | | `opensuse-leap-15.6` | `opensuse` | 10 GiB | qcow2 | | `arch` | `arch` | 10 GiB | qcow2 | | `talos-v1.13.9` | — none, Talos has no SSH | 10 GiB | raw (xz) | Every instance type's root disk is at least 20 GiB ([Instance types](/docs/compute/ec2/ec2-api-instance-types)), so the minimum root disk column never blocks a launch. It matters if you shrink a root disk through a block device mapping. There is no Windows image, and no licence to sell one. ## The login user is not `root` Cloud images disable password login and root SSH, and create one unprivileged user with your key in `~/.ssh/authorized_keys` and passwordless `sudo`. Which user that is, is set by the distribution — not by us. ``` ssh debian@ # not root@, not ubuntu@ ``` `DescribeImages` states the login user in the image `description` for exactly this reason. ## Talos `talos-v1.13.9` is Talos Linux, and it is not like the others: - **No SSH, no shell, no login user.** It is an API-driven Kubernetes OS; you talk to it with `talosctl`. - Shipped as a raw image, xz-compressed, from the Talos image factory. - The schematic id in its URL names its extensions: this build **carries the Tailscale extension**, which is why it is the build on the shelf. - It expects machine configuration through the cloud-init user-data channel, the same channel `RunInstances --user-data` writes. Pick it when you are running Kubernetes and you want the node OS to be immutable. Pick Debian when you want a machine you can SSH into. ## Names roll An image name points at the newest build we have imported. Re-running the importer moves the name to the newer upstream build **and the `ami-` id changes**. | If you | Then | | --- | --- | | Write `ubuntu-24.04` in Terraform via a `data "aws_ami"` lookup | You get whatever we imported most recently, which is usually what you want | | Write a literal `ami-…` id | You get that exact build, until it is deregistered | | Need immutability across months | **Record the image checksum**, not the id | This is deliberate: the vendors publish rolling `latest` URLs (Debian, Ubuntu, Rocky, Alma, openSUSE, Arch all do), and re-running the play is how the offer stays patched. The one exception is Fedora, which publishes no stable `latest` URL — `fedora-42` is pinned and bumped by hand. > [!warning] > > **Image immutability is an open gap.** our internal readiness list gap 7 asks for a > recorded fingerprint per image id, or dated ids. Until it is done, an > `ami-` id you stored is not guaranteed to mean the same bytes it meant last > month. Pin by checksum if that matters to you. ## What is not offered - **No public or shared images from other accounts.** Every image `DescribeImages` returns is one we published. - **No customer-built images.** `CreateImage`, `RegisterImage`, `CopyImage` and `DeregisterImage` are `OPEN (Michael)` on the [actions page](/docs/compute/ec2/api-reference-actions) — they are the difference between customers using our images and customers building their own, and that decision has not been taken. - **No marketplace, no paid AMIs, no bring-your-own-licence.** ## Finding one ``` aws --endpoint-url https://ec2.shelfcs.com --region hel1 \ ec2 describe-images --filters Name=name,Values=debian-13 ``` ```hcl data "aws_ami" "debian" { most_recent = true filter { name = "name" values = ["debian-13"] } } ``` Through the cloud platform directly, the same images are in the image service at `https://image.shelfcs.com`: ``` cloud image list ``` ## Go further - [Image and catalogue actions](/docs/compute/ec2/ec2-api-images-and-catalogue) — `DescribeImages` itself - [Instance types](/docs/compute/ec2/ec2-api-instance-types) - [Key pairs](/docs/compute/ec2/ec2-api-key-pairs) - [RunInstances](/docs/compute/ec2/ec2-api-run-instances) --- ## Making requests Source: https://shelfcs.com/docs/compute/ec2/api-reference-making-requests Request structure, signing, retries, eventual consistency and idempotency ## Objective **Read this guide to construct a request by hand, or to understand what your SDK does on your behalf.** ## Requirements - An access key id and secret access key ## Instructions ### Request structure A request is a `POST` to the service endpoint with a form-encoded body: ``` POST / HTTP/1.1 Host: ec2.hel1.shelfcloud.com Content-Type: application/x-www-form-urlencoded; charset=utf-8 X-Amz-Date: 20260903T142207Z Authorization: AWS4-HMAC-SHA256 Credential=AKIA.../20260903/hel1/ec2/aws4_request, SignedHeaders=content-type;host;x-amz-date, Signature=... Action=DescribeInstances&Version=2016-11-15 ``` `GET` with a query string is also accepted, and is what presigned requests use. ### Signing The `Credential` parameter is the access key id, a slash, then the credential scope: ``` AKIAIOSFODNN7EXAMPLE/20260903/hel1/ec2/aws4_request ``` The access key id is sent in the clear and is **not** part of the credential scope. The scope is the four trailing elements. The hashed payload is always the final line of the canonical request, including for requests with no body, which sign the hash of the empty string. `SignedHeaders` must contain `host`, `x-amz-date`, and every `x-amz-*` header present on the request, plus `content-type` when a body is sent. Headers not listed in `SignedHeaders` are ignored by the server rather than acted on. ### Clock skew **PROPOSED:** 15 minutes. A request whose `X-Amz-Date` is outside that window is rejected with `RequestExpired` and HTTP 400. Our server time is in the `Date` header of every response, including error responses, so that a client can detect its own skew. The AWS SDKs correct their clocks automatically from this, and only when the status and error code match what they expect — which is why the status codes on the error page are not negotiable. ### Retries Retry on: - HTTP 500, 502, 503, 504 - `RequestLimitExceeded` - `InternalError` Retry with exponential backoff and jitter. ### Eventual consistency A resource that has just been created may not be visible to an immediately following call. `DescribeInstances` may not return an instance that `RunInstances` has just returned an id for, and a security group referenced seconds after creation may return `InvalidGroup.NotFound`. This is inherited from Amazon EC2 and existing tooling already accounts for it: the Terraform AWS provider retries `NotFound` errors on recently created resources for exactly this reason. Therefore, and against the general rule that a 4xx is not retryable, **do retry a `NotFound` error for a resource you created in the last few seconds.** **PROPOSED:** a created resource is visible to all readers within 10 seconds. ### Idempotency `RunInstances` accepts a `ClientToken`. A repeated request with the same token returns the original result rather than launching a second instance. If the token is reused with different parameters, the request is refused with `IdempotentParameterMismatch`. Actions that do not accept a `ClientToken` in the Amazon EC2 model do not accept one here either. The SDKs validate parameters against their own model before sending, so a token added to an action AWS does not model would be rejected by the client, before it ever reached us. For those actions, retry safety comes from the action's own semantics: `CreateTags` sets tags to a value and repeating it is harmless; `DeleteVolume` on an already-deleted volume returns `InvalidVolume.NotFound`, which the caller treats as success. ### Pagination See [Pagination](/docs/compute/ec2/api-reference-pagination). ## Go further - [Common errors](/docs/compute/ec2/api-reference-common-errors) - [Authentication conventions](/docs/platform/conventions/authentication) --- ## Networking Source: https://shelfcs.com/docs/compute/ec2/ec2-api-networking The network, subnet and router every customer project gets at onboarding, the default VPC on top of them, security groups, and where addresses come from. ## Objective The [actions page](/docs/compute/ec2/api-reference-actions) once listed networking as `OPEN` and put VPCs and subnets under "not implemented". **That is out of date on both counts** — there is per-tenant networking, created for you, and a default VPC on top of it. ## What you get, and when `POST /v1/onboard` creates four the networking service objects in your project, in order, and looks each one up by name before creating it — so a retry after a failure finishes the job rather than making a second of everything: | Object | Name | Detail | | --- | --- | --- | | Network | `default` | Yours, in your project | | Subnet | `default` | `10.200.0.0/24`, DNS `9.9.9.9` and `149.112.112.112` | | Router | `default` | External gateway on the `guests` network | | Router interface | — | The router attached to your subnet | You do not create these and you do not need to. They exist before your first launch. ## Every project gets the same private range `10.200.0.0/24`. All of them. That is deliberate, not a collision waiting to happen: **projects are isolated by their own router**, so two customers' `10.200.0.5` are on different networks behind different routers and never meet. The address you get is not a hint about how many customers there are, and it is stable for you. It also means you cannot pick your own CIDR. There is one subnet, it is a `/24`, and it gives you 253 usable addresses. ## Why the per-project network exists at all `ec2-api` places a machine only on a **non-external network the calling project owns**. Without a network of your own there is nothing for a launch to attach to — so `ensureNetwork` is not a nicety, it is what makes `RunInstances` work. The `guests` network is the **external** network: the gateway your router points at, and the pool floating addresses come from. It is not where your instances live. ## EC2-Classic is off, so you have a default VPC ``` disable_ec2_classic = True ``` The reason is Terraform. Amazon retired EC2-Classic and the AWS provider assumes it is gone: the provider **revokes the default egress rule of every security group it creates**, and this API refuses that on a classic group. With classic disabled, each project gets a default VPC on first use, and its security groups, instances and addresses are VPC ones — the shape the provider expects. Consequences: - A launch that names no subnet lands in your default VPC. - Security groups have egress rules as well as ingress rules; the default group allows all egress until something revokes it. - `aws_security_group` applies without the classic-mode error. ## Security groups | Action | Status | | --- | --- | | `CreateSecurityGroup` | Available | | `DescribeSecurityGroups` | Available | | `AuthorizeSecurityGroupIngress` | Available | | `AuthorizeSecurityGroupEgress` | Available | | `RevokeSecurityGroupIngress` | Available | | `RevokeSecurityGroupEgress` | Available | | `DeleteSecurityGroup` | Available | They map onto the networking service security groups: stateful, deny-by-default on ingress, and evaluated as a union when several are attached to one port — the same semantics as EC2. A group that references another security group as its source works, because the networking service supports remote group ids. A group written for AWS behaves the same way here. ## Addresses Every instance gets a private address on your own subnet. **A public address is not automatic.** `DescribeAddresses` is available. Floating addresses are allocated from the `guests` external network your router already points at. **OPEN (Michael):** whether `AllocateAddress` and `AssociateAddress` are offered through the EC2 API. Through the cloud platform directly the networking service floating-IP calls are the ones to use — see [Account developer guide](/docs/account/guide/developer-guide): ``` cloud floating ip create guests cloud server add floating ip ``` Outbound traffic works without any of this: your router does NAT. ## Isolation, stated precisely Customer projects are separated at layer 3 by their own network and router. Another customer is not on your broadcast domain and cannot reach your instances by private address. What that does **not** give you: - **No microsegmentation inside your own project.** Every instance of yours is on one flat `/24` and can reach every other. Security groups are the tool for separating your own tiers, and they are the only tool. - **No network ACLs**, no subnet-level policy, no route control. - **No multiple subnets**, so you cannot put a database on a subnet with no route out. > [!primary] > > `the infrastructure repository` gap 8 says every VM shares one > bridge and can reach every other VM, and calls per-tenant networking a hard > blocker. That note predates `ensureNetwork` and describes the platform team's > own shared `guests` network, not a customer project. The per-project network, > subnet and router described above are what a customer account actually gets. ## Not implemented Named so their absence is deliberate. Each returns `InvalidAction`: Customer-managed VPCs and additional subnets, route tables, internet gateways, NAT gateways, VPC peering, transit gateways, VPN connections, VPC endpoints, network ACLs, Elastic Network Interfaces as first-class resources, flow logs, and IPv6. There is also **no load balancer** — Octavia is not deployed and there is no `elasticloadbalancing` endpoint. ## Go further - [EC2 API actions](/docs/compute/ec2/api-reference-actions) - [Account API](/docs/account/cloud/storefront-api-reference) — where the network is created - [Account developer guide](/docs/account/guide/developer-guide) — the networking service directly - [Instance metadata and user data](/docs/compute/ec2/ec2-instance-metadata-and-user-data) - [Service endpoints](/docs/platform/endpoints/service-endpoints) --- ## Pagination Source: https://shelfcs.com/docs/compute/ec2/api-reference-pagination Which actions paginate, how tokens behave, and why an empty page is not the end ## Objective **Read this guide before writing a loop over any describe action.** The most common integration bug against this API is stopping on an empty page. ## Requirements - None ## Instructions ### Which actions paginate Only the actions whose Amazon EC2 model declares `MaxResults` and `NextToken`. `DescribeInstances`, `DescribeVolumes`, `DescribeSnapshots` and `DescribeImages` paginate. `DescribeRegions` and `DescribeAvailabilityZones` do not, and never return a token. This is not a stylistic decision. The SDKs validate parameters against their own model before sending, and their paginators only exist for actions the model marks as paginated. A token returned by an action AWS models as unpaginated is silently discarded by the client, and the caller sees a truncated list with no error at all. ### Parameters | Parameter | Type | Notes | | --- | --- | --- | | `MaxResults` | integer | Items per page. **PROPOSED:** 5 to 1000, default 1000 | | `NextToken` | string | From a previous response | `MaxResults` cannot be combined with an explicit list of resource ids. That combination returns `InvalidParameterCombination`. ### The token `NextToken` is opaque. Do not parse it, construct it, or store it beyond the sequence it belongs to. A response that omits `NextToken` is the last page. **That is the only signal that a listing has ended.** A page may contain fewer items than `MaxResults`, or none at all, and still carry a `NextToken`. This happens whenever filters remove every item from a page, and it is normal. A loop that stops on an empty page silently misses results. ```python token = None while True: kwargs = {"MaxResults": 1000} if token: kwargs["NextToken"] = token page = ec2.describe_instances(**kwargs) handle(page["Reservations"]) token = page.get("NextToken") if not token: # not: if not page["Reservations"] break ``` **PROPOSED:** a token is valid for 24 hours, after which `InvalidPaginationToken`. ### Filters Filters are applied per page, not before pagination. This matches Amazon EC2 and is why an empty page with a token occurs. ### Ordering and consistency **PROPOSED:** results are ordered by creation time, oldest first, stable across pages. A listing is not a snapshot. Items created during a pagination sequence may or may not appear. Items deleted during one may still appear, and a later read of one may return `.NotFound`. Reconcile after the listing completes rather than assuming it gave you a consistent view. ## Go further - [Making requests](/docs/compute/ec2/api-reference-making-requests) - [Actions](/docs/compute/ec2/api-reference-actions) --- ## RunInstances Source: https://shelfcs.com/docs/compute/ec2/ec2-api-run-instances Launch instances — every parameter, what is accepted, what is refused, the response, the errors, and worked examples in CLI, Terraform and boto3. Launches one or more instances. This is the only action that creates a machine. `RunInstances` is **idempotent through `ClientToken`**. Always send one. ## Request | Parameter | Type | Required | Notes | | --- | --- | --- | --- | | `ImageId` | string | Yes | `ami-…`. See [Machine images](/docs/compute/ec2/ec2-api-machine-images) | | `MinCount` | integer | Yes | ≥ 1 | | `MaxCount` | integer | Yes | ≥ 1, and ≥ `MinCount` | | `InstanceType` | string | No | Our name or an AWS alias. Defaults to `cd-standard-2-2` | | `KeyName` | string | No | A key pair you have imported. Without it you cannot log in | | `UserData` | string | No | Base64. **≤ 16384 bytes decoded** | | `SecurityGroupId.N` | list | No | Defaults to the default group of your VPC | | `SubnetId` | string | No | Defaults to your own subnet | | `PrivateIpAddress` | string | No | Must be inside your subnet | | `TagSpecification.N` | list | No | `ResourceType` must be `instance` or `volume` | | `BlockDeviceMapping.N` | list | No | Root device size only | | `ClientToken` | string | No | ≤ 64 ASCII characters. **Send one** | | `Placement.AvailabilityZone` | string | No | `hel1-a` | | `MetadataOptions` | structure | No | Accepted only at the platform defaults — see below | | `DryRun` | boolean | No | Evaluate permissions and do nothing else | ### `MinCount` and `MaxCount` The call launches as many as it can between the two. If it cannot reach `MinCount` it launches **nothing** and returns `InsufficientInstanceCapacity` — it does not partially succeed. For one machine, set both to 1. ### `ClientToken` Two calls with the same token return the same instances rather than launching a second set. A token is remembered long enough to cover a retry after a lost response, which is exactly the case it exists for. Without one, a network timeout between your client and us can leave you paying for a machine you do not know about. Every SDK will generate one if you ask; the CLI takes `--client-token`. ### `UserData` Base64, at most **16384 bytes decoded** — the Amazon limit, enforced here even though the platform underneath would accept more, so that a payload which works on AWS works here unchanged. Never logged: the parameter is marked sensitive. Delivery, the config drive, and why a `running` instance may still be unconfigured: [Instance metadata and user data](/docs/compute/ec2/ec2-instance-metadata-and-user-data). ### `MetadataOptions` Accepted only when every member given equals the platform default: ``` HttpTokens=optional HttpEndpoint=enabled HttpPutResponseHopLimit=1 HttpProtocolIpv6=disabled InstanceMetadataTags=disabled ``` Anything else is `UnsupportedOperation`. In particular `HttpTokens=required` (IMDSv2) is **refused rather than silently downgraded**, because the metadata service has no session-token mode and pretending otherwise would leave you believing in a defence you do not have. Terraform sends this block only when you configure `metadata_options`. ### `BlockDeviceMapping.N` Supported for the **root device only**, and only to change its size. Additional volumes are created and attached separately — see [Volume actions](/docs/compute/ec2/ec2-api-volumes). A root device size below the image's minimum is `InvalidParameterValue`. ### Parameters that are refused Refused with `InvalidParameterValue` or `UnsupportedOperation`, never ignored: `InstanceMarketOptions` (spot), `CapacityReservationSpecification`, `Placement.GroupName`, `Placement.HostId` and `Placement.Tenancy`, `LaunchTemplate`, `IamInstanceProfile`, `HibernationOptions`, `ElasticGpuSpecification`, `ElasticInferenceAccelerators`, `EnclaveOptions`, `CpuOptions`, `CreditSpecification`, `Ipv6AddressCount` and `Ipv6Addresses`, `NetworkInterface.N` beyond a single default interface. Refusing rather than ignoring is deliberate. A module that asks for an IAM instance profile fails at apply rather than launching a machine that silently has no credentials. ## Response A reservation containing one entry per instance launched: | Element | Notes | | --- | --- | | `reservationId` | `r-…` | | `instancesSet` | One `instance` per machine | | `instanceId` | `i-…` | | `imageId`, `instanceType`, `keyName` | As launched. **`instanceType` is our name**, even if you passed an AWS alias | | `instanceState` | `pending` at this point | | `privateIpAddress` | Present once the address is assigned | | `placement.availabilityZone` | `hel1-a` | | `amiLaunchIndex` | Position within the reservation | | `clientToken` | If you sent one | | `tagSet` | Tags applied at launch | | `blockDeviceMapping` | The root device | | `groupSet` | Security groups | > [!primary] > > **The response names our instance type, not the alias you launched with.** > Launch `m5.large` and `DescribeInstances` reports `cd-standard-2-8`. Terraform > records that in state, so a configuration written with the AWS name shows a > diff on every plan. Write our name, or `ignore_changes = [instance_type]`. Tags applied through `TagSpecification` are applied **as part of the launch**, not afterwards, so a machine never exists untagged. That is what makes tag-based reconciliation safe after a failed retry. ## Errors | Code | Status | Cause | | --- | --- | --- | | `InvalidAMIID.NotFound` | 400 | No such image, or not one we publish | | `InvalidParameterValue` | 400 | Unknown instance type, `MinCount`/`MaxCount` < 1, user data not base64 or over the limit, client token over 64 characters, a bad tag | | `InvalidParameterCombination` | 400 | `MaxCount` below `MinCount` | | `InvalidKeyPair.NotFound` | 400 | No such key pair in your account | | `InvalidGroup.NotFound` | 400 | No such security group | | `InstanceLimitExceeded` | 400 | Your vCPU or memory quota | | `InsufficientInstanceCapacity` | 500 | The hardware cannot fit it now. **Not a quota problem** | | `UnsupportedOperation` | 400 | A parameter that exists in the model but not here | | `Unsupported` | 400 | A combination that cannot be honoured | `InsufficientInstanceCapacity` is the one to handle differently: retrying the same shape or asking for more quota will not help. Take a smaller shape, or wait. `GET /v1/catalog` reports whether a shape is in stock right now — [Catalog and availability](/docs/billing/pricing/catalog-and-availability). ## Examples ### AWS CLI ```bash aws --endpoint-url https://ec2.shelfcs.com --region hel1 \ ec2 run-instances \ --image-id ami-… \ --instance-type cd-standard-2-4 \ --key-name mykey \ --count 1 \ --client-token "$(uuidgen)" \ --user-data file://cloud-init.yaml \ --tag-specifications 'ResourceType=instance,Tags=[{Key=Name,Value=web}]' ``` `--count 1` sets both `MinCount` and `MaxCount`. ### Terraform ```hcl resource "aws_instance" "web" { ami = data.aws_ami.debian.id instance_type = "cd-standard-2-4" # our name, not m5.large key_name = aws_key_pair.mine.key_name user_data = file("cloud-init.yaml") tags = { Name = "web" } } ``` The provider generates a client token for you. ### boto3 ```python res = ec2.run_instances( ImageId=image_id, InstanceType="cd-standard-2-4", KeyName="mykey", MinCount=1, MaxCount=1, ClientToken=str(uuid.uuid4()), TagSpecifications=[{ "ResourceType": "instance", "Tags": [{"Key": "Name", "Value": "web"}], }], ) iid = res["Instances"][0]["InstanceId"] ec2.get_waiter("instance_running").wait(InstanceIds=[iid]) ``` ## After it returns `pending` becomes `running` in well under a minute. That means the machine booted — **not** that your configuration applied. See [Instance metadata and user data](/docs/compute/ec2/ec2-instance-metadata-and-user-data). Then connect as the image's login user, never `root` — [Machine images](/docs/compute/ec2/ec2-api-machine-images). ## Go further - [DescribeInstances](/docs/compute/ec2/ec2-api-describe-instances) - [TerminateInstances](/docs/compute/ec2/ec2-api-terminate-instances) - [Instance types](/docs/compute/ec2/ec2-api-instance-types) - [Compute user guide](/docs/compute/guide/user-guide) - [Idempotency](/docs/platform/conventions/idempotency) --- ## Snapshot actions Source: https://shelfcs.com/docs/compute/ec2/ec2-api-snapshots CreateSnapshot, DescribeSnapshots, DeleteSnapshot — what a snapshot really guarantees, how to restore from one, and the capacity it quietly consumes. A snapshot is a point-in-time copy of a volume. On a platform with **no provider backups**, snapshots are the only recovery mechanism you have, so it is worth knowing exactly what they do and do not protect against. ## Permissions | Action | Resource scope | | --- | --- | | `ec2:CreateSnapshot` | Both `volume/` and `snapshot/*` | | `ec2:DescribeSnapshots` | Requires `"Resource": "*"` | | `ec2:DeleteSnapshot` | `snapshot/` | Every action accepts `DryRun`. ## CreateSnapshot | Parameter | Type | Required | Notes | | --- | --- | --- | --- | | `VolumeId` | string | Yes | Attached or detached, either works | | `Description` | string | No | Up to 255 characters | | `TagSpecification.N` | list | No | `ResourceType` must be `snapshot` | | `DryRun` | boolean | No | | **Response:** `snapshotId`, `volumeId`, `status` (`pending`), `startTime`, `progress`, `volumeSize`, `description`, `tagSet`, `ownerId`, `encrypted` (`false`). The call returns as soon as the snapshot is registered. Copying continues afterwards and `status` moves `pending` → `completed`. **A snapshot cannot create a volume until it is `completed`.** ```python snap = ec2.create_snapshot(VolumeId=vol, Description="before upgrade") ec2.get_waiter("snapshot_completed").wait(SnapshotIds=[snap["SnapshotId"]]) ``` Not idempotent, and there is no `ClientToken` for it. A retry after a lost response creates a second snapshot occupying the same space again. Tag at creation and reconcile by tag. ## What a snapshot captures Whatever was on the disk at the instant it was taken — **including a filesystem part-way through a write**. | Workload | Do this first | | --- | --- | | A database | Stop it, or use its own dump or hot-backup mechanism | | A filesystem that can freeze | `fsfreeze -f /data`, snapshot, `fsfreeze -u /data` | | Anything else | Unmount, or accept crash consistency | Crash-consistent means the snapshot is exactly what the disk would look like after a power cut. Most filesystems recover from that; most databases do not guarantee it. ```bash sudo fsfreeze -f /data sc ec2 create-snapshot --volume-id vol-… --description "nightly" sudo fsfreeze -u /data ``` Keep the frozen window short — writes block for its duration. > [!warning] > > **A snapshot is stored on the same hardware as the volume.** It protects you > from your own mistakes — a bad deploy, a wrong `rm`, a migration you want to > undo — not from the failure of the machine holding both. > > It is not a backup, we do not describe it as one, and anything you cannot lose > must be copied off this platform. ## Restoring Restore by creating a **new volume** from the snapshot, then attaching it: ```bash sc ec2 create-volume \ --availability-zone hel1-a \ --snapshot-id snap-… \ --volume-type standard sc ec2 attach-volume --volume-id vol-NEW --instance-id i-… --device /dev/sdg ``` - `Size` may be omitted to take the snapshot's size, or given to create a **larger** volume. It may never be smaller. - The restored volume is a **new volume with a new id**. Nothing is restored in place, and the original volume is untouched. - If you grew the volume during restore, grow the filesystem inside the guest afterwards — see [Volume actions](/docs/compute/ec2/ec2-api-volumes). Restoring beside the original rather than over it is the safer habit: mount the restored volume somewhere else, check it, then swap. ## DescribeSnapshots | Parameter | Type | Notes | | --- | --- | --- | | `SnapshotId.N` | list | | | `Filter.N` | list | `status`, `volume-id`, `volume-size`, `start-time`, `progress`, `description`, `tag:`, `tag-key` | | `OwnerId.N` | list | Only your own account resolves | | `MaxResults` | integer | 5–1000 | | `NextToken` | string | | Paginated. Returns only snapshots your account owns — **there are no public or shared snapshots**, and none from other accounts. Filters naming things this platform never reports — encryption, storage tier, restore state — return `InvalidParameterValue` rather than matching nothing. ## DeleteSnapshot | Parameter | Type | Required | | --- | --- | --- | | `SnapshotId` | string | Yes | | `DryRun` | boolean | No | Permanent, with no recycle bin. A volume already created from the snapshot is **unaffected** — it is a volume in its own right from the moment it is created, not a reference to the snapshot. ## They consume the pool, and nothing expires them > [!warning] > > **Snapshots occupy the same storage pool as every volume and every instance > root disk.** The pool is finite and shared. > > Nothing expires a snapshot. There is no lifecycle policy, no retention rule > and no scheduling. A nightly snapshot taken by a cron job you forgot about > accumulates until something fails to create. > > When the pool is full, `CreateVolume` returns `InsufficientVolumeCapacity` — > and that failure lands on whoever asks next, not necessarily on the account > whose snapshots filled it. Practical rule: **if you take snapshots on a schedule, delete them on a schedule too.** Nothing else will. ```bash # snapshots older than 30 days, oldest first sc ec2 describe-snapshots --owner-ids self \ --query 'sort_by(Snapshots,&StartTime)[?StartTime<`2026-08-05`].[SnapshotId,StartTime,VolumeSize]' \ --output table ``` ## Billing **Snapshots are not rated today.** Nothing charges for the storage they occupy. That is a gap in the rating configuration, not a discount — see [How prices are set](/docs/billing/pricing/how-prices-are-set). It will change, and when it does, forgotten snapshots become an invoice line. The capacity they consume is real now regardless of what they cost. ## Errors | Code | Status | Cause | | --- | --- | --- | | `InvalidSnapshot.NotFound` | 400 | No such snapshot, or another account's | | `InvalidSnapshot.InUse` | 400 | A volume is still being created from it | | `InvalidVolume.NotFound` | 400 | No such volume to snapshot | | `IncorrectState` | 400 | The volume is not in a state that can be snapshotted | | `InsufficientVolumeCapacity` | 500 | The pool cannot hold it | | `InvalidParameterCombination` | 400 | Restore size smaller than the snapshot | ## Not implemented `CopySnapshot`, `ModifySnapshotAttribute`, `ResetSnapshotAttribute`, `CreateSnapshots` (multi-volume), snapshot sharing, fast snapshot restore, archive tiers, and the EBS direct APIs return `InvalidAction`. There is **no cross-region or off-platform copy**. Getting data out is your job, through the filesystem. ## Go further - [Volume actions](/docs/compute/ec2/ec2-api-volumes) - [Volume types](/docs/storage/block/volume-types) — the pool they share - [Storage user guide](/docs/storage/guide/user-guide) - [Storage developer guide](/docs/storage/guide/developer-guide) --- ## StartInstances, StopInstances, RebootInstances Source: https://shelfcs.com/docs/compute/ec2/ec2-api-start-stop-instances Change the run state of an instance, and what each transition costs. ## Objective Three actions that change whether an instance is running. They differ in what survives, what is charged, and whether the instance may move to another machine. ## Permissions | Action | Resource scope | ARN shape | | --- | --- | --- | | `ec2:StartInstances` | Resource-scoped | `instance/` | | `ec2:StopInstances` | Resource-scoped | `instance/` | | `ec2:RebootInstances` | Resource-scoped | `instance/` | ## Instructions ### StopInstances Shuts the operating system down and releases CPU and memory. The root volume is kept. | Parameter | Type | Required | Notes | | --- | --- | --- | --- | | `InstanceId.N` | list | Yes | | | `Force` | boolean | No | Stop without a clean shutdown | | `DryRun` | boolean | No | | `Force` skips the guest shutdown. It is the equivalent of removing power, and it risks filesystem corruption in the same way. Use it when a guest has stopped responding, not as a default. **Response:** `instancesSet`, each with `instanceId`, `previousState` and `currentState`. ```xml b1e2c3d4-5678-90ab-cdef-1234567890ab i-0a1b2c3d4e5f60718 64stopping 16running ``` **Billing:** rating counts a **running** instance, by its shape, per hour — so the charge ends when the instance leaves `running`. Its root disk is not billed separately, because storage is not rated at all today. It does still **occupy the pool**, so a stopped instance costs you nothing and costs the platform capacity. See [Billing user guide](/docs/billing/guide/user-guide). **The public address is not retained.** A stopped instance that is started again may receive a different address. Stopping an already-stopped instance succeeds and reports its current state. ### StartInstances Starts a stopped instance. | Parameter | Type | Required | | --- | --- | --- | | `InstanceId.N` | list | Yes | | `DryRun` | boolean | No | The instance returns to `pending`, then `running`. **It may be placed on a different physical machine**, which is why a start can fail for capacity when the original launch did not. **Errors** | Code | Status | Cause | | --- | --- | --- | | `IncorrectInstanceState` | 400 | The instance is not `stopped` | | `InsufficientInstanceCapacity` | 500 | No capacity for that type now | | `InvalidInstanceID.NotFound` | 400 | No such instance | `InsufficientInstanceCapacity` on a start is worth calling out: an instance you stopped is not capacity held for you. Stopping releases it, and starting competes for it again. ### RebootInstances Restarts the guest operating system in place. | Parameter | Type | Required | | --- | --- | --- | | `InstanceId.N` | list | Yes | | `DryRun` | boolean | No | The instance does not leave `running`, does not change machine, keeps its address, and continues to be charged. The response carries no instance state, only a request id and `return`. A reboot is not a stop-then-start, and cannot be used to move an instance to a different machine or to change its address. ### Which to use | You want to | Use | | --- | --- | | Restart the OS | `RebootInstances` | | Stop paying for compute, keep the disk | `StopInstances` | | Stop paying entirely, lose the disk | `TerminateInstances` | | Move to another machine | Stop, then start | ## Go further - [TerminateInstances](/docs/compute/ec2/ec2-api-terminate-instances) - Instance lifecycle --- ## Tag actions Source: https://shelfcs.com/docs/compute/ec2/ec2-api-tags CreateTags, DeleteTags, DescribeTags, and tagging at creation. ## Objective Tags are key-value pairs attached to resources. They organise resources, filter `Describe` calls, and appear on usage records. ## Permissions | Action | Resource scope | | --- | --- | | `ec2:CreateTags` | The resource being tagged | | `ec2:DeleteTags` | The resource being untagged | | `ec2:DescribeTags` | Requires `"Resource": "*"` | `ec2:CreateTags` is also required by any create action that carries a `TagSpecification`. A policy that allows `ec2:RunInstances` but not `ec2:CreateTags` denies a launch that tags the instance. ## Instructions ### Tag constraints **PROPOSED**, matching Amazon EC2 so that existing tooling's client-side validation agrees with ours: | Constraint | Value | | --- | --- | | Tags per resource | 50 | | Key length | 1 to 128 characters | | Value length | 0 to 256 characters | | Character set | Letters, digits, spaces, and `+ - = . _ : / @` | | Case | Keys and values are case-sensitive | | Reserved prefix | `shelf:` is reserved and cannot be set by a customer | A value may be empty. A key may not. > [!warning] > > Tags are visible to anyone who can read the resource, they appear in usage > records, and they may appear on invoices. Nothing sensitive belongs in a tag. ### Tagging at creation Prefer this to tagging afterwards: ``` TagSpecification.1.ResourceType=instance TagSpecification.1.Tag.1.Key=Name TagSpecification.1.Tag.1.Value=web-1 ``` The resource is never untagged, even briefly. A separate `CreateTags` call after a create leaves a window in which the resource exists without its tags — and if the second call fails, an untagged resource that nothing is tracking. ### CreateTags | Parameter | Type | Required | | --- | --- | --- | | `ResourceId.N` | list | Yes | | `Tag.N.Key`, `Tag.N.Value` | list | Yes | | `DryRun` | boolean | No | Sets tags to the values given. An existing key is overwritten; keys not mentioned are left alone. **Idempotent by nature**, which is why it carries no `ClientToken` and needs none: repeating the call with the same values changes nothing. Resources of different types may be tagged in one call. It is all-or-nothing: if any resource id is invalid, none are tagged. **Response:** `return` only. No tag data is echoed. ### DeleteTags | Parameter | Type | Required | Notes | | --- | --- | --- | --- | | `ResourceId.N` | list | Yes | | | `Tag.N.Key` | list | No | Omit to delete all tags | | `Tag.N.Value` | list | No | See below | The value parameter has behaviour worth reading twice: - **Key given, no value** — the tag is deleted whatever its value. - **Key and value given** — deleted only if the value matches. A mismatch is not an error; nothing is deleted and the call succeeds. - **Key given, value given as empty** — deleted only if the value is empty. Deleting a tag that does not exist succeeds. ### DescribeTags | Parameter | Type | Notes | | --- | --- | --- | | `Filter.N` | list | `key`, `value`, `resource-id`, `resource-type` | | `MaxResults` | integer | 5 to 1000 | | `NextToken` | string | | **Paginated.** Returns one entry per tag per resource, so a hundred resources with five tags each is five hundred entries. **Response:** each entry has `resourceId`, `resourceType`, `key`, `value`. ### Filtering other calls by tag Every `Describe` action accepts tag filters: ```bash shelf --profile shelf ec2 describe-instances \ --filters Name=tag:Environment,Values=production ``` | Filter | Matches | | --- | --- | | `tag:` | Resources whose tag `` has one of the given values | | `tag-key` | Resources carrying that key, whatever the value | This is the mechanism most automation depends on, and the reason to tag at creation: a resource that missed its tags is invisible to every query that selects by tag. ## Go further - [RunInstances](/docs/compute/ec2/ec2-api-run-instances) - [DescribeInstances](/docs/compute/ec2/ec2-api-describe-instances) --- ## TerminateInstances Source: https://shelfcs.com/docs/compute/ec2/ec2-api-terminate-instances Destroy instances permanently, what goes with them, and what the response reports. ## Objective `TerminateInstances` destroys instances and their root volumes. It cannot be undone. ## Requirements - Credentials configured ## Permissions | Action | Resource scope | ARN shape | | --- | --- | --- | | `ec2:TerminateInstances` | Resource-scoped | `arn:aws-shelf:ec2:::instance/` | A policy may name specific instances, or use a wildcard. A policy granting `ec2:TerminateInstances` on `"Resource": "*"` permits terminating everything in the account. ## Instructions ### Request parameters | Parameter | Type | Required | Notes | | --- | --- | --- | --- | | `InstanceId.N` | list | Yes | One or more instance ids | | `DryRun` | boolean | No | Check permissions without acting | ### What is destroyed - The instance - Its root volume, and everything written to it - Its public address ### What survives - Volumes attached separately, which are detached and kept - Snapshots - Key pairs > [!warning] > > There is no recovery. We hold no copy of the root volume, support cannot > restore one, and there is no console through which to rescue anything. ### Behaviour The call returns immediately with each instance in `shutting-down`. Destruction completes shortly afterwards and the state becomes `terminated`. Terminating an already-terminated instance succeeds and reports its current state. Terminating an instance that has never existed returns `InvalidInstanceID.NotFound`. That distinction is load-bearing: infrastructure tooling polls for a not-found error to confirm a destroy completed, and treats not-found during a refresh as "deleted out of band". A server that reported success for resources it had never heard of would leave that tooling unable to tell "destroyed" from "never existed". ### Response | Element | Notes | | --- | --- | | `requestId` | | | `instancesSet` | One entry per instance | | `instancesSet.N.instanceId` | | | `instancesSet.N.previousState` | The state before this call | | `instancesSet.N.currentState` | Usually `shutting-down` | `previousState` is how a caller distinguishes "I terminated this" from "it was already gone": a `previousState` of `terminated` means the instance was already destroyed before this call. ```xml b1e2c3d4-5678-90ab-cdef-1234567890ab i-0a1b2c3d4e5f60718 32shutting-down 16running ``` ### Partial failure If any id in the list is invalid, the whole call fails with `InvalidInstanceID.NotFound` and **nothing is terminated**. Validate first, or terminate one at a time when the list is assembled programmatically. ### Errors | Code | Status | Cause | | --- | --- | --- | | `InvalidInstanceID.NotFound` | 400 | An instance does not exist | | `InvalidInstanceID.Malformed` | 400 | An id is not well-formed | | `UnauthorizedOperation` | 403 | Policy denies termination | | `DryRunOperation` | 400 | `DryRun` was set and the call would have succeeded | | `OperationNotPermitted` | 400 | Termination protection is enabled | **OPEN (Michael):** whether termination protection exists in v1. It is the only guard between a wrong instance id and permanent data loss, and its absence should be a decision rather than an omission. ### Testing a policy without destroying anything ```bash aws --profile shelf ec2 terminate-instances \ --instance-ids i-0a1b2c3d4e5f60718 --dry-run ``` `DryRunOperation` means the call would have succeeded. `UnauthorizedOperation` means policy would have denied it. Nothing is terminated either way. ### Billing Charging for the instance stops when it enters `shutting-down` — rating counts a running instance, and it is no longer one. Separately attached volumes are **not** billed today, because storage is not rated; they do survive the instance and keep occupying the pool until you delete them. Unbilled is not free — see [Billing user guide](/docs/billing/guide/user-guide). ## Go further - [Start and stop instances](/docs/compute/ec2/ec2-api-start-stop-instances) - [Volume actions](/docs/compute/ec2/ec2-api-volumes) - [Volume types](/docs/storage/block/volume-types) - [RunInstances](/docs/compute/ec2/ec2-api-run-instances) --- ## Volume actions Source: https://shelfcs.com/docs/compute/ec2/ec2-api-volumes CreateVolume, DescribeVolumes, AttachVolume, DetachVolume, ModifyVolume, DeleteVolume, DescribeVolumeStatus — over the block-storage service, with the storage this hardware actually delivers. Terms used below: **the pool** is the single storage pool every volume, and every instance root disk, is carved from. **The record** is the layer's own map from a `vol-`, `snap-` or `i-` identifier to the underlying resource. **Assumption: RunInstances §9.4 Option A** — the `vol-` ↔ the block-storage service’s own UUID binding lives in a store owned by the layer, because the block-storage service’s own `metadata` on a volume is customer-writable and therefore cannot hold an identifier the layer trusts. Under Option B (the block-storage service volume-image metadata property) every id resolution below becomes `GET /volumes/detail?metadata={"vol-id":…}`, one call per id, and a customer who edits that metadata orphans their own volume. **Nothing on this page is enforced that the catalog does not enforce.** Where a figure is advertised but not yet limited on the hypervisor, it is marked `NOT ENFORCED` and named in section 9. That is the opposite of how the rest of this platform works, and it is a defect, not a design. ## 1. Conformance `Implemented with differences`. | # | AWS | This layer | Why | What breaks for a client assuming AWS | | --- | --- | --- | --- | --- | | D1 | Volume types `gp2`, `gp3`, `io1`, `io2`, `st1`, `sc1`, `standard`. | One type. `standard` is accepted and returned; every other name is rejected with `InvalidParameterValue`. | The hardware has one class of disk — 7200 rpm SATA. `standard` is Amazon's own name for magnetic storage, so the alias is honest; `gp3` and `io2` would be a promise of SSD latency this hardware cannot make. | Terraform `aws_ebs_volume` with `type = "gp3"` errors instead of silently getting a slower disk. This is deliberate — see section 9.1. | | D2 | `Iops` and `Throughput` are settable on `CreateVolume` and `ModifyVolume` for the provisioned types. | Both parameters are rejected with `InvalidParameterValue`. The performance of a volume is fixed by its type and published in [Volume types](/docs/storage/block/volume-types). | There is one type and it is not provisionable. Accepting the parameter and ignoring it would let a customer believe they had bought IOPS. | A module that always sets `iops` must set it conditionally on type, which is what the AWS provider already requires. | | D3 | Size 1 GiB to 16 TiB depending on type. | 1 GiB to 1000 GiB (`volume_types[0].min_gib`, `max_gib`). Outside that: `InvalidParameterValue`. | The whole pool is 1400 GiB and it also holds every instance root disk. | A request for a 2 TiB volume errors rather than being silently truncated. | | D4 | Multi-attach on `io1`/`io2`. | A volume attaches to one instance at a time. `MultiAttachEnabled` is rejected. | Not implemented, and not safe on a single-node LVM backend without a cluster filesystem. | `aws_ebs_volume.multi_attach_enabled` errors. | | D5 | `Encrypted`, `KmsKeyId`. | Rejected with `InvalidParameterValue`. There is no key management service. | No KMS exists on this platform. Returning `encrypted: false` for a request that asked for `true` would be a lie a compliance audit would later find. | A module that sets `encrypted = true` by policy fails loudly at plan time rather than shipping unencrypted storage. | | D6 | `DescribeVolumes` `MaxResults` 5–500. | 5–500, default 1000 when absent, with a `nextToken` when more exist. | Conventions, "Pagination". | None. | | D7 | A deleted volume "might appear" briefly. | `DeleteVolume` returns after the block-storage service accepts the delete; the volume then reports `deleting` until the block-storage service finishes, and afterwards `InvalidVolume.NotFound`. There is no synthesised `deleted` window. | the block-storage service's delete is asynchronous and the object disappears at the end of it; unlike a terminated instance, no client polls a volume *to* a terminal state — the AWS provider's `aws_ebs_volume` delete waits for not-found. | None. | | D8 | Another account's volume id behaves as not-found. | `InvalidVolume.NotFound`, byte-identical to an id that never existed. | Conventions, "Accounts and tenancy", Visibility. | None. | | D9 | Volume IOPS and throughput are enforced per volume. | **`NOT ENFORCED`.** The catalog declares 48 IOPS and 11 MB/s per volume; no a block-storage QoS specification or front-end throttle applies it today. A single busy volume can take the whole pool's throughput. | `bootstrap/roles/openstack/tasks/resources.yml` creates the compute service flavors from the catalog and nothing else; there is no `volume_type`/`qos_spec` task. | A noisy neighbour. This breaks the platform's own first rule — every capability figure is enforced, not advertised — and it blocks the first customer. Section 9.2. | ## 2. Endpoint and credentials Volume actions are EC2 actions. They go to the EC2 endpoint, not to the block-storage service: ``` https://ec2.shelfcs.com ``` Signed with SigV4 using an access key pair, verified by the identity service's the identity service’s signature-verification call. See [Service endpoints](/docs/platform/endpoints/service-endpoints) for the full host list and [Authentication](/docs/platform/conventions/authentication) for the signing rules. The same volumes are reachable through the block-storage service's own API at `https://volume.shelfcs.com` — see [Storage developer guide](/docs/storage/guide/developer-guide). The two views are of one object; a volume created through EC2 is visible, resizable and deletable through the block-storage service and vice versa. Only the identifier differs (`vol-…` against a UUID). ## 3. Permissions | Action | Resource scope | | --- | --- | | `ec2:CreateVolume` | `volume/*` | | `ec2:DescribeVolumes` | Requires `"Resource": "*"` | | `ec2:DescribeVolumeStatus` | Requires `"Resource": "*"` | | `ec2:AttachVolume` | Both `volume/` and `instance/` | | `ec2:DetachVolume` | Both `volume/` and `instance/` | | `ec2:ModifyVolume` | `volume/` | | `ec2:DeleteVolume` | `volume/` | `AttachVolume` and `DetachVolume` evaluate against **both** resources. A policy naming only the volume denies the call. This is Amazon's rule and it is kept, because a policy written for AWS must not become more permissive here. Every action accepts `DryRun`. Permitted: `DryRunOperation`. Denied: `UnauthorizedOperation` (HTTP 403). ## 4. CreateVolume | Parameter | Model member | Type | Required | Constraints | Status | | --- | --- | --- | --- | --- | --- | | `AvailabilityZone` | `AvailabilityZone` | string | Yes | `hel1-a`. No default. Unknown zone: `InvalidParameterValue`. | `supported` | | `Size` | `Size` | integer | Conditional | GiB, 1–1000. Required unless `SnapshotId` is given. | `supported` | | `VolumeType` | `VolumeType` | string | No | `standard` only. Defaults to `standard`. | `supported` | | `SnapshotId` | `SnapshotId` | string | No | `snap-…`. Restores from a snapshot. | `supported` | | `TagSpecification.N` | `TagSpecifications` | list | No | `ResourceType` must be `volume`. | `supported` | | `DryRun` | `DryRun` | boolean | No | | `supported` | | `Iops`, `Throughput` | | integer | No | | `rejected with InvalidParameterValue` (D2) | | `Encrypted`, `KmsKeyId` | | | No | | `rejected with InvalidParameterValue` (D5) | | `MultiAttachEnabled` | | boolean | No | | `rejected with InvalidParameterValue` (D4) | | `OutpostArn` | | string | No | | `rejected with InvalidParameterValue` | `AvailabilityZone` is required and has no default. A volume attaches only to an instance in its own zone, and there is no cross-zone attach. There is one zone today; ask `DescribeAvailabilityZones` for it rather than writing `hel1-a` into your code, so that the day there are two you do not have to. With `SnapshotId`, `Size` may be omitted to take the snapshot's size, or given to create a larger volume. It may never be smaller than the snapshot. ### Mapping `POST /v3/{project_id}/volumes`, body `volume`: | EC2 | the block-storage service | | --- | --- | | `Size` | `size` | | `AvailabilityZone` | `availability_zone` | | `VolumeType` (`standard`) | `volume_type` `hdd` | | `SnapshotId` (resolved) | `snapshot_id` | | `TagSpecification.N` | `metadata`, key-prefixed so customer metadata and tags cannot collide | The `vol-` identifier is minted by the layer and written to the record against the UUID the block-storage service returns, before the response is sent. A crash between the block-storage service's 201 and that write leaves an orphaned the block-storage service volume that the EC2 API cannot see; it is still billed. This is the same failure the missing `ClientToken` causes below, and it is reconciled by the same sweep — section 9.3. ### Response `volumeId`, `size`, `volumeType` (`standard`), `status` (`creating`), `availabilityZone`, `createTime`, `snapshotId` where applicable, `tagSet`, `encrypted` (`false`), `multiAttachEnabled` (`false`). `volumeType` is returned as `standard`, which is both the value accepted on input and the value a client's model expects. Our own name for it, `hdd`, is never on the wire. ### Errors | Code | Status | Cause | | --- | --- | --- | | `InvalidParameterValue` | 400 | Size out of 1–1000, unknown type, unknown zone, or a rejected parameter (D1, D2, D3, D4, D5) | | `InvalidParameterCombination` | 400 | Neither `Size` nor `SnapshotId`; or `Size` smaller than the snapshot | | `InvalidSnapshot.NotFound` | 400 | No such snapshot, or it belongs to another account | | `VolumeLimitExceeded` | 400 | The project's the block-storage service quota would be exceeded | | `InsufficientVolumeCapacity` | 500 | The pool cannot fit it (section 6) | ### Not idempotent The Amazon EC2 model declares no `ClientToken` for `CreateVolume`, so none can be sent — an SDK rejects the parameter before the request leaves the machine. A retry after a lost response creates a second volume, and bills for it. Tag volumes at creation and reconcile by tag if you retry automatically. Terraform does this for you: `aws_ebs_volume` records the id in state before it retries anything. ## 5. DescribeVolumes | Parameter | Type | Constraints | | --- | --- | --- | | `VolumeId.N` | list | `^vol-([0-9a-f]{8}\|[0-9a-f]{17})$`. Anything else: `InvalidVolumeID.Malformed` | | `Filter.N` | list | See below | | `MaxResults` | integer | 5–500. Cannot be combined with `VolumeId.N` | | `NextToken` | string | Opaque. Cannot be combined with `VolumeId.N` | | `DryRun` | boolean | | Paginated. An empty page may still carry a `NextToken` — filters are applied after the page is assembled (conventions, "Pagination", Filters), so a page can filter down to nothing and still not be the last. ### Filters Supported — a filter is supported when the element it matches is emitted: | Filter | Matched against | | --- | --- | | `status` | `creating`, `available`, `in-use`, `deleting`, `error` | | `attachment.instance-id` | The instance a volume is attached to | | `attachment.device` | The requested device name | | `attachment.status` | `attaching`, `attached`, `detaching` | | `attachment.delete-on-termination` | | | `availability-zone` | | | `size` | GiB, exact | | `snapshot-id` | | | `volume-type` | Always `standard` | | `create-time` | Wildcards `*` and `?` allowed | | `volume-id` | | | `tag:`, `tag-key` | | Rejected with `InvalidParameterValue`: `encrypted`, `multi-attach-enabled`, `fast-restored`, `throughput`, `iops`, `outpost-arn`. They describe elements this layer never emits, and matching nothing silently is the failure the pagination convention forbids. ### Response For each volume: `volumeId`, `size`, `snapshotId`, `availabilityZone`, `status`, `createTime`, `attachmentSet`, `volumeType`, `encrypted` (`false`), `multiAttachEnabled` (`false`), `tagSet`. Each `attachmentSet` item: `volumeId`, `instanceId`, `device`, `status`, `attachTime`, `deleteOnTermination`. ### Volume state Computed from the block-storage service `status`: | the block-storage service | EC2 | | --- | --- | | `creating`, `downloading` | `creating` | | `available` | `available` | | `attaching`, `in-use`, `detaching` | `in-use` | | `deleting` | `deleting` | | `error`, `error_deleting`, `error_extending`, `error_restoring` | `error` | | `reserved` | `available` — the block-storage service reserves a volume between the attach request and the attach; EC2 has no such state | | `maintenance`, `backing-up`, `restoring-backup`, `retyping` | `NOT MAPPED` — section 9.4 | ## 6. Capacity, and what "thin" means here Every volume and every instance root disk comes out of one pool: | Figure | Value | Source | | --- | --- | --- | | Pool | `default` | | Sellable | 1400 GiB | | Reserved for instance root disks | 1200 GiB thin LV `nova-instances` | | Volume GiB already sold | Tracked internally; free space plus that figure must still cover the catalog | Root volumes are thin-provisioned, but the **sum of what has been sold is capped**, so the pool cannot fill silently behind a customer who is within their quota. A `CreateVolume` that would take the pool past that cap fails with `InsufficientVolumeCapacity` rather than succeeding into an over-committed pool. > [!warning] > > One box, one pool, one disk class. There is no replication of a volume across > hosts, and as of today the pool is not mirrored across disks. See > our internal readiness list gap 1. Treat a volume as durable against a process > crash, not against a disk failure, and take snapshots. ## 7. AttachVolume | Parameter | Type | Required | Notes | | --- | --- | --- | --- | | `VolumeId` | string | Yes | Must be `available` | | `InstanceId` | string | Yes | Must be in the same zone, and `running` or `stopped` | | `Device` | string | Yes | The device name inside the guest | | `DryRun` | boolean | No | | Mapped to the compute service `POST /servers/{server_id}/os-volume_attachments` with `volumeId` and `device` — the compute service, not the block-storage service, because the compute service owns the hypervisor side of the attach and calling the block-storage service's `os-attach` directly would leave the compute service's own view of the instance wrong. The device name is a request, not a guarantee: the guest kernel decides what the device is actually called, and `/dev/sdf` may arrive as `/dev/vdb`. Identify volumes inside the instance by filesystem UUID or label — `blkid`, then `UUID=…` in `/etc/fstab` — or a reboot that renumbers devices mounts the wrong one. Attaching does not partition, format or mount anything. That happens inside the instance, and we do not do it for you. **Errors** | Code | Status | Cause | | --- | --- | --- | | `VolumeInUse` | 400 | Already attached to an instance | | `InvalidVolume.ZoneMismatch` | 400 | Volume and instance are in different zones | | `IncorrectState` | 400 | The volume is not `available`, or the instance is not in a state that can take an attach | | `InvalidParameterValue` | 400 | The device name is already in use on that instance | | `InvalidInstanceID.NotFound` | 400 | No such instance, or another account's | | `AttachmentLimitExceeded` | 400 | Per-instance attachment limit reached | A volume attaches to one instance at a time. There is no multi-attach (D4). `deleteOnTermination` is `false` for every volume attached by this action. A volume attached after launch survives the instance. Only a root volume created by `RunInstances` from a block device mapping carries `deleteOnTermination: true`. ## 8. DetachVolume | Parameter | Type | Required | Notes | | --- | --- | --- | --- | | `VolumeId` | string | Yes | | | `InstanceId` | string | No | Verified if given; mismatch is `InvalidParameterValue` | | `Device` | string | No | Verified if given | | `Force` | boolean | No | Detach without the guest releasing it | | `DryRun` | boolean | No | | Mapped to the compute service `DELETE /servers/{server_id}/os-volume_attachments/{volume_id}`. With `Force`, to the block-storage service `os-force_detach`, which drops the attachment record whether or not the hypervisor agreed. > [!warning] > > Unmount the filesystem inside the instance **before** detaching. Detaching a > mounted, written filesystem is the equivalent of pulling the disk out, and > loses whatever was in flight. > > `Force` detaches regardless. It is for an unresponsive instance, and it risks > filesystem corruption. It does not stop the instance first. Detaching does not stop billing. The volume still exists and is still charged until it is deleted (section 10). ## 9. ModifyVolume | Parameter | Type | Required | Notes | | --- | --- | --- | --- | | `VolumeId` | string | Yes | | | `Size` | integer | Yes | Larger than the current size, at most 1000 | | `DryRun` | boolean | No | | | `VolumeType`, `Iops`, `Throughput`, `MultiAttachEnabled` | | No | `rejected with InvalidParameterValue` | Mapped to the block-storage service `POST /volumes/{volume_id}/action`, `os-extend`. Volumes grow and never shrink. A `Size` at or below the current size is `InvalidParameterValue`, not a no-op — a silent no-op would let a shrink attempt look like it worked. **Response:** `volumeModification` with `volumeId`, `modificationState` (`modifying`, then `optimizing`, then `completed`), `originalSize`, `targetSize`, `startTime`. `DescribeVolumesModifications` is **not implemented**; poll `DescribeVolumes` for the new `size` instead. A client that waits on `DescribeVolumesModifications` receives `InvalidAction` — section 12. Growing the volume does **not** grow the filesystem on it. Inside the instance, extend the partition and then the filesystem — `growpart /dev/vdb 1`, then `resize2fs` or `xfs_growfs`. Until you do, the extra space is invisible to the guest and you are paying for it. Extending an attached volume is supported by the block-storage service for a volume in `in-use`, but the guest only sees the new size after a rescan (`echo 1 > /sys/class/block/vdb/device/rescan`) or a reboot. ## 10. DeleteVolume | Parameter | Type | Required | | --- | --- | --- | | `VolumeId` | string | Yes | | `DryRun` | boolean | No | The volume must be `available`, not attached. Deleting an attached volume is `VolumeInUse`. > [!warning] > > Deletion destroys the data. We hold no copy, and there is no recycle bin. > Snapshots taken before deletion survive and are billed separately. Deleting an already-deleted volume returns `InvalidVolume.NotFound`. Callers retrying a delete treat that as success — the AWS provider does. **Billing:** the charge ends when the volume is destroyed, not when it is detached. See [How prices are set](/docs/billing/pricing/how-prices-are-set) for what a volume-GiB-hour currently costs, which is a number that does not yet exist — attached-storage rating is not implemented, so volumes are today **unbilled**. That is a revenue defect, not a customer discount, and it will change. ## 11. DescribeVolumeStatus | Parameter | Type | Notes | | --- | --- | --- | | `VolumeId.N` | list | | | `Filter.N` | list | `volume-status.status`, `availability-zone`, `volume-status.details-name`, `volume-status.details-status` | | `MaxResults` | integer | 5–1000 | | `NextToken` | string | | Paginated. **Response:** `volumeId`, `availabilityZone`, `volumeStatus.status` (`ok`, `impaired`, `insufficient-data`), `volumeStatus.details`, `actionsSet` (always empty), `eventsSet` (always empty). `volumeStatus.status` is derived from the block-storage service `status` alone: `error*` states map to `impaired`, everything else to `ok`. There is no per-volume health check behind it, so `ok` means "the block-storage service has not recorded an error", not "we have verified this disk is healthy". SMART monitoring runs at the host level (`smartd`), not per volume, and does not feed this field. `eventsSet` is always empty: there are no scheduled-maintenance events, because there is no maintenance scheduling system. A client that watches this field for retirement notices will never see one — that is a real gap, not a promise that nothing will ever fail. ## 12. Not implemented Named so that their absence is deliberate. Each returns `InvalidAction`: `DescribeVolumesModifications`, `EnableVolumeIO`, `ModifyVolumeAttribute`, `DescribeVolumeAttribute`, `CreateVolumePermission` and the whole `*VolumeAttribute` family, `AttachVolume` with `MultiAttachEnabled`, fast snapshot restore, EBS direct APIs (`ListChangedBlocks`, `GetSnapshotBlock`), Elastic Volumes' online type change, and every provisioned-performance action. ## 13. Conformance evidence ### AWS CLI ``` aws --endpoint-url https://ec2.shelfcs.com --region hel1 \ ec2 create-volume --availability-zone hel1-a --size 20 --volume-type standard \ --tag-specifications 'ResourceType=volume,Tags=[{Key=Name,Value=data}]' aws --endpoint-url https://ec2.shelfcs.com --region hel1 \ ec2 attach-volume --volume-id vol-… --instance-id i-… --device /dev/sdf aws --endpoint-url https://ec2.shelfcs.com --region hel1 \ ec2 describe-volumes --filters Name=attachment.instance-id,Values=i-… ``` `--region hel1` must match the region in the credential scope, or every call returns `SignatureDoesNotMatch` with no diagnosable cause. ### Terraform ```hcl resource "aws_ebs_volume" "data" { availability_zone = "hel1-a" size = 20 type = "standard" tags = { Name = "data" } } resource "aws_volume_attachment" "data" { device_name = "/dev/sdf" volume_id = aws_ebs_volume.data.id instance_id = aws_instance.app.id } ``` `type = "standard"` is required: the provider's default is `gp2`, which this layer rejects (D1). Omitting `type` therefore fails at apply, not at plan. `aws_volume_attachment` sets `force_detach = false` by default; leave it there. `skip_destroy = true` is the safe setting when the volume outlives the instance. ### Inside the instance ``` lsblk mkfs.ext4 /dev/vdb blkid /dev/vdb # take the UUID echo 'UUID= /data ext4 defaults,nofail 0 2' >> /etc/fstab mount -a ``` `nofail` matters: without it an instance whose volume is detached fails to boot into anything you can SSH to, and there is no serial console on the EC2 API — you would need the [web console](/docs/platform/console/web-console-and-vnc). ## 14. Open decisions Each is marked `OPEN (Michael)` and none has been filled with a plausible value. 1. **Whether `gp3` is accepted as an alias for `standard`.** Accepting it makes every stock Terraform module apply unchanged, which is the whole thesis of an AWS-shaped API; it also means a customer who asked for 3000 IOPS gets 48 and is told nothing. The current answer is to reject, and to lose the modules. 2. **Volume QoS (D9).** The catalog's 48 IOPS / 11 MB/s is advertised and not enforced. Enforcing it needs a block-storage QoS specification associated with the `hdd` volume type and a task in `resources.yml` to create both — today that file creates flavors only. Until then the platform's own first rule is broken. 3. **Reconciling orphaned the block-storage service volumes** created when the layer crashed between the block-storage service's 201 and the record write (section 4). A periodic sweep of the block-storage service volumes with no `vol-` in the record, and what it does with them — delete, or adopt. 4. **the block-storage service states with no EC2 mapping** (section 5): `maintenance`, `backing-up`, `restoring-backup`, `retyping`. An operator action can put a volume into one and the interim rule reports `in-use`, which is a guess. 5. **Whether volumes are billed, and at what rate.** Section 10. The catalog's pricing rule prices the AWS-equivalent instance type only; there is no storage price rule, so `cd-storage-4-16-1000` was removed from the catalog for exactly this reason and attached volumes are free by accident. 6. **Snapshot storage cost and retention.** A snapshot survives its volume and nothing expires it. ## Go further - [Snapshot actions](/docs/compute/ec2/ec2-api-snapshots) - [Volume types](/docs/storage/block/volume-types) - [Storage developer guide](/docs/storage/guide/developer-guide) - [Service endpoints](/docs/platform/endpoints/service-endpoints) - [RunInstances](/docs/compute/ec2/ec2-api-run-instances) --- # Compute — GUIDE ## Compute cheat sheet Source: https://shelfcs.com/docs/compute/guide/cheat-sheet Every command you need on one page — setup, launch, connect, storage, teardown, and the four errors you will actually hit. ## Setup ```bash export AWS_ACCESS_KEY_ID=… # console → Access keys, or POST /v1/access-keys export AWS_SECRET_ACCESS_KEY=… alias sc='aws --endpoint-url https://ec2.shelfcs.com --region hel1' ``` ```ini # ~/.aws/config — so you stop typing the flags [profile shelf] region = hel1 endpoint_url = https://ec2.shelfcs.com ``` Region is **`hel1`**. Zone is **`hel1-a`**. ## Look around ```bash sc ec2 describe-instance-types # shapes sc ec2 describe-images # operating systems sc ec2 describe-availability-zones # hel1-a sc ec2 describe-instances # what you have sc ec2 describe-volumes sc ec2 describe-security-groups ``` ```bash curl -s https://storefront.job-rss-processor.workers.dev/v1/catalog | jq # prices + live stock ``` ## Keys ```bash sc ec2 import-key-pair --key-name mykey \ --public-key-material fileb://~/.ssh/id_ed25519.pub sc ec2 describe-key-pairs sc ec2 delete-key-pair --key-name mykey ``` Import your own. `create-key-pair` returns the private key once and never again. ## Launch ```bash sc ec2 run-instances \ --image-id ami-… \ --instance-type cd-standard-2-4 \ --key-name mykey \ --count 1 \ --client-token "$(uuidgen)" \ --tag-specifications 'ResourceType=instance,Tags=[{Key=Name,Value=web}]' ``` ```bash sc ec2 run-instances … --user-data file://cloud-init.yaml # ≤16384 bytes decoded ``` Always send `--client-token`: it makes a retry safe. ## Connect ```bash sc ec2 describe-instances --instance-ids i-… \ --query 'Reservations[].Instances[].[InstanceId,PrivateIpAddress,PublicIpAddress,State.Name]' \ --output table ssh debian@
# NOT root@ — the user is set by the image ``` | Image | User | | --- | --- | | `debian-12`, `debian-13` | `debian` | | `ubuntu-22.04`, `ubuntu-24.04` | `ubuntu` | | `rocky-9` | `rocky` | | `alma-9` | `almalinux` | | `fedora-42` | `fedora` | | `opensuse-leap-15.6` | `opensuse` | | `arch` | `arch` | | `talos-v1.13.9` | none — no SSH | Cannot get in? The [VNC console](/docs/platform/console/web-console-and-vnc) works before the network does. ## Lifecycle ```bash sc ec2 stop-instances --instance-ids i-… # keeps the disk and the address sc ec2 start-instances --instance-ids i-… sc ec2 reboot-instances --instance-ids i-… sc ec2 terminate-instances --instance-ids i-… # final; root disk goes with it ``` ## Storage ```bash sc ec2 create-volume --availability-zone hel1-a --size 20 --volume-type standard sc ec2 attach-volume --volume-id vol-… --instance-id i-… --device /dev/sdf sc ec2 modify-volume --volume-id vol-… --size 40 # grows only, never shrinks sc ec2 detach-volume --volume-id vol-… # unmount inside FIRST sc ec2 delete-volume --volume-id vol-… # destroys the data sc ec2 create-snapshot --volume-id vol-… --description "before upgrade" ``` `--volume-type standard` is required — the AWS default is `gp2`, which is rejected. There is one type, on 7200 rpm SATA. Inside the guest: ```bash lsblk mkfs.ext4 /dev/vdb blkid /dev/vdb echo 'UUID= /data ext4 defaults,nofail 0 2' >> /etc/fstab # nofail matters mount -a growpart /dev/vdb 1 && resize2fs /dev/vdb1 # after modify-volume ``` ## Security groups ```bash sc ec2 create-security-group --group-name web --description "http" sc ec2 authorize-security-group-ingress --group-id sg-… \ --protocol tcp --port 443 --cidr 0.0.0.0/0 sc ec2 revoke-security-group-ingress --group-id sg-… \ --protocol tcp --port 443 --cidr 0.0.0.0/0 ``` They are the only boundary inside your account — everything you run is on one flat `10.200.0.0/24`. ## Terraform ```hcl provider "aws" { region = "hel1" skip_credentials_validation = true skip_requesting_account_id = true skip_metadata_api_check = true endpoints { ec2 = "https://ec2.shelfcs.com" } } resource "aws_instance" "web" { ami = data.aws_ami.debian.id instance_type = "cd-standard-2-4" # write OUR name, not m5.large key_name = aws_key_pair.mine.key_name user_data = file("cloud-init.yaml") } resource "aws_ebs_volume" "data" { availability_zone = "hel1-a" size = 20 type = "standard" # the gp2 default is rejected } ``` ## Test a call without doing it ```bash sc ec2 terminate-instances --instance-ids i-… --dry-run ``` `DryRunOperation` = permitted. `UnauthorizedOperation` = not. ## The four errors you will actually hit | Error | Fix | | --- | --- | | `SignatureDoesNotMatch` | Region. It is `hel1` | | `InsufficientInstanceCapacity` | Box is full. Smaller shape, or wait — quota will not help | | `InstanceLimitExceeded` | Your quota. Starts at 4 vCPU / 8 GB | | `InvalidParameterValue` on `type=gp3` | One volume type, `standard` | ## Debugging a machine that booted but does not work ```bash cloud-init status --long sudo journalctl -u cloud-final -b sudo cat /var/log/cloud-init-output.log # configuration is delivered on a drive, not over the network mount /dev/disk/by-label/config-2 /mnt && cat /mnt/openstack/latest/user_data ``` `runcmd` swallows failures — a `running` instance is not a configured one. ## Go further - [User guide](/docs/compute/guide/user-guide) - [Developer guide](/docs/compute/guide/developer-guide) - [API reference](/docs/compute/ec2/api-reference-welcome) --- ## Compute developer guide Source: https://shelfcs.com/docs/compute/guide/developer-guide Working with the compute API itself — endpoint, credentials, signing a request, pagination, idempotency, error handling, and code examples in four languages. ## What this is Developer-focused information about using the compute API: how a request is made, how it is authenticated and signed, what comes back, and what to do when what comes back is an error. For what the service *is* and how to get your first machine running, read the [user guide](/docs/compute/guide/user-guide). For every operation and its parameters, the [API reference](/docs/compute/ec2/api-reference-welcome). ## Endpoint and region ``` https://ec2.shelfcs.com ``` Region `hel1`. The region is not in the hostname but it **is** in the credential scope of every signature, so every client must be told it explicitly. Getting it wrong returns `SignatureDoesNotMatch` and nothing in the message says why. There is no endpoint discovery service. Hosts are constructed from the table in [Service endpoints](/docs/platform/endpoints/service-endpoints). HTTPS only. HTTP is refused, never redirected — a redirect would invite a client to send a signed request unencrypted first. ## Credentials An **access key pair**, minted from the console or `POST /v1/access-keys`. The signature is verified by the identity service, so the key and your cloud password are two faces of one identity: revoking the user revokes both. Full detail, including the one way these do not behave like AWS: [Access keys](/docs/account/cloud/access-keys). There are **no temporary credentials**, no role assumption, and no instance profile — a machine cannot authenticate as itself. See [Identity, as deployed](/docs/identity/keystone/identity-as-deployed). ## Making a request The API speaks the EC2 query protocol: `POST`, form-encoded, `Action` and `Version` in the body, signed with **Signature Version 4**. You do not have to build that by hand, and should not. Every AWS SDK, the AWS CLI and the Terraform AWS provider already speak it; point them at the endpoint. ``` POST / HTTP/1.1 Host: ec2.shelfcs.com Content-Type: application/x-www-form-urlencoded Authorization: AWS4-HMAC-SHA256 Credential=…/20260904/hel1/ec2/aws4_request, … Action=DescribeInstances&Version=2016-11-15 ``` Details of the wire format, headers and the request-id contract: [Making requests](/docs/compute/ec2/api-reference-making-requests) and [Requests and responses](/docs/platform/conventions/requests-and-responses). ## Signing SigV4, with the credential scope `/hel1/ec2/aws4_request`. The service name is `ec2` and the region is `hel1`. This is why a single global `endpoint_url` override in a client configuration is wrong: it sends every service's calls to one host, and those calls are signed for the service they were meant for. See [Authentication](/docs/platform/conventions/authentication). ## Code examples ### AWS CLI ```bash export AWS_ACCESS_KEY_ID=… export AWS_SECRET_ACCESS_KEY=… aws --endpoint-url https://ec2.shelfcs.com --region hel1 ec2 describe-instances ``` Put it in a profile so you stop typing the flags: ```ini # ~/.aws/config [profile shelf] region = hel1 endpoint_url = https://ec2.shelfcs.com ``` ### Python — boto3 ```python import boto3 ec2 = boto3.client( "ec2", endpoint_url="https://ec2.shelfcs.com", region_name="hel1", ) for page in ec2.get_paginator("describe_instances").paginate(): for res in page["Reservations"]: for i in res["Instances"]: print(i["InstanceId"], i["InstanceType"], i["State"]["Name"]) ``` Use the paginator rather than reading one page — the API returns a `nextToken` and a client that ignores it silently sees a partial answer. ### JavaScript — AWS SDK v3 ```js import { EC2Client, DescribeInstancesCommand } from "@aws-sdk/client-ec2"; const ec2 = new EC2Client({ endpoint: "https://ec2.shelfcs.com", region: "hel1", }); const out = await ec2.send(new DescribeInstancesCommand({ MaxResults: 100 })); ``` ### Go — aws-sdk-go-v2 ```go cfg, _ := config.LoadDefaultConfig(ctx, config.WithRegion("hel1"), config.WithBaseEndpoint("https://ec2.shelfcs.com"), ) svc := ec2.NewFromConfig(cfg) out, err := svc.DescribeInstances(ctx, &ec2.DescribeInstancesInput{}) ``` ### Terraform ```hcl provider "aws" { region = "hel1" access_key = var.access_key secret_key = var.secret_key skip_credentials_validation = true skip_requesting_account_id = true skip_metadata_api_check = true endpoints { ec2 = "https://ec2.shelfcs.com" } } ``` The three `skip_*` lines are required: the provider otherwise calls STS to validate credentials and find an account id, and there is no STS here. ## Pagination Paginated actions return a `nextToken`. Keep asking until it is absent — do not assume one page is the whole answer. Three actions are deliberately **not** paginated, because their Amazon models are not: `DescribeRegions`, `DescribeAvailabilityZones`, `DescribeAccountAttributes`. Every SDK builds no paginator for them and would discard a token, so returning one would silently truncate a location list. They return everything in one response, always. Filters are applied **after** a page is assembled, so an empty page can still carry a token. See [Pagination](/docs/compute/ec2/api-reference-pagination). ## Idempotency and retries | Action | Retry behaviour | | --- | --- | | `RunInstances` | Idempotent via `ClientToken`. Always send one | | `TerminateInstances`, `StartInstances`, `StopInstances` | Naturally idempotent | | `CreateTags` | Sets values; repeating changes nothing | | `CreateVolume` | **Not idempotent, and no `ClientToken` exists** — a retry after a lost response creates a second volume, and bills for it | | `CreateKeyPair`, `ImportKeyPair` | Not idempotent | Tag resources at creation and reconcile by tag if you retry automatically. See [Idempotency](/docs/platform/conventions/idempotency). ## Error handling Errors are the EC2 shape: an HTTP status, an error `Code`, a `Message`, and a `RequestID`. Quote the request id in any support case. | Class | Examples | Do | | --- | --- | --- | | Malformed input | `InvalidParameterValue`, `InvalidParameterCombination` | Fix the call. Retrying will not help | | Not found / not yours | `InvalidInstanceID.NotFound`, `InvalidVolume.NotFound` | Another account's id is byte-identical to one that never existed. Treat as not-found | | State | `IncorrectState`, `VolumeInUse` | Wait for the resource to reach the state you need, then retry | | Quota | `InstanceLimitExceeded`, `VolumeLimitExceeded` | Ask for more quota | | Capacity | `InsufficientInstanceCapacity`, `InsufficientVolumeCapacity` | The box is full. A quota increase will not help; take a smaller shape or wait | | Auth | `SignatureDoesNotMatch`, `UnauthorizedOperation` | Check the region first. It is `hel1` | | Unsupported | `InvalidAction`, `UnsupportedOperation` | The call or that parameter is not implemented. See the action list | `DryRun` on any action evaluates permissions and does nothing else: `DryRunOperation` if permitted, `UnauthorizedOperation` if not. It is how you test before running something destructive. Full tables: [Common errors](/docs/compute/ec2/api-reference-common-errors) and [Errors](/docs/platform/conventions/errors). ## Differences worth knowing before you write code - **`InvalidAction` rather than approximation.** An action not on the [action list](/docs/compute/ec2/api-reference-actions) errors instead of doing something nearly right. - **Rejected parameters, not ignored ones.** `Encrypted`, `Iops`, `Throughput`, `MultiAttachEnabled` and `gp3` are refused rather than silently dropped, so a module that asks for encryption fails at plan rather than shipping unencrypted. - **Responses name our instance types**, not the AWS alias you launched with — write ours in Terraform or you get a perpetual diff. - **No IMDSv2.** `HttpTokens=required` is refused, not downgraded. - **No rate limiting exists yet.** Do not read that as permission; it will arrive, and code that hammers the API will notice. ## Go further - [User guide](/docs/compute/guide/user-guide) — concepts and getting started - [API reference](/docs/compute/ec2/api-reference-welcome) — every operation - [Cheat sheet](/docs/compute/guide/cheat-sheet) - [Authentication](/docs/platform/conventions/authentication) - [Access keys](/docs/account/cloud/access-keys) --- ## Compute user guide Source: https://shelfcs.com/docs/compute/guide/user-guide Getting started with compute — what a machine is made of, how to go from an account to something running, and what you have to do yourself. ## What this is Compute here is **one virtual machine on dedicated cores**, launched through an API shaped like Amazon EC2, so the tooling you already have works against it unchanged. It is not a managed service. Nothing restarts your application, nothing fails it over, nothing patches it. You get a machine and the machine is yours. **Read this guide first.** Then the [developer guide](/docs/compute/guide/developer-guide) for how to sign and send a request, the [API reference](/docs/compute/ec2/api-reference-welcome) for every operation, and the [cheat sheet](/docs/compute/guide/cheat-sheet) when you already know what you want. ## What a machine is made of Five things, and you choose four of them: | Piece | You choose | Where it is documented | | --- | --- | --- | | **Shape** — cores and memory | Yes | [Instance types](/docs/compute/ec2/ec2-api-instance-types) | | **Image** — the operating system | Yes | [Machine images](/docs/compute/ec2/ec2-api-machine-images) | | **Key pair** — how you get in | Yes | [Key pairs](/docs/compute/ec2/ec2-api-key-pairs) | | **User data** — how it configures itself | Optional | [Metadata and user data](/docs/compute/ec2/ec2-instance-metadata-and-user-data) | | **Network** — where it lives | No, you already have one | [Networking](/docs/compute/ec2/ec2-api-networking) | The network is the one you do not choose: your account came with a network, a subnet and a router, created at signup. There is nothing to build before your first launch. ## The lifecycle ``` RunInstances → pending → running → StopInstances → stopped → StartInstances → running ↘ TerminateInstances → shutting-down → terminated ``` - **`pending` → `running`** is under a minute. Your key is injected by cloud-init during it. - **`stopped`** still holds your root disk and still holds your private address. A stopped machine is not a deleted one, and its storage still occupies the pool. - **`terminated`** is final. The root disk goes with it. Attached volumes do not — see below. ## Storage, and what survives This is the distinction that costs people data: | | Root disk | Attached volume | | --- | --- | --- | | Comes from | The shape, 20–160 GB | `CreateVolume`, 1–1000 GB | | Survives terminate | **No** | **Yes** | | Growable | No | Yes, never shrinks | | Cost | In the instance price | Not billed today | The disk underneath both is **7200 rpm SATA**, quoted at 48 IOPS and 11 MB/s per volume — a figure divided by the most machines the box can hold, so it is one every customer can have at once. Read [Volume types](/docs/storage/block/volume-types) before you put a database on it. ## What you have to do yourself Stated plainly, because none of it is done for you: - **Backups.** There are none. Snapshots are yours to take and they live in the same pool as the volume. - **Redundancy.** One machine, one box, no zone to fail over to. - **Patching, monitoring, alerting, log shipping.** Yours. - **Separating your own tiers.** Everything you run is on one flat `/24`. Security groups are the only boundary inside your account. - **Credentials for anything the machine calls.** There are no instance roles; a key on disk is the only option. ## Getting from nothing to running 1. **Sign up and onboard.** `POST /v1/onboard` creates your project, your cloud user, your network and your Stripe customer — [Account API](/docs/account/cloud/storefront-api-reference). 2. **Mint an access key.** The console's Access keys page, or `POST /v1/access-keys` — [Access keys](/docs/account/cloud/access-keys). 3. **Point your tooling at the endpoint.** `https://ec2.shelfcs.com`, region `hel1`. 4. **Import a key pair**, choose an image and a shape, launch. 5. **SSH as the image's login user** — `debian`, `ubuntu`, `rocky`… not `root`. Every command for those steps is on the [cheat sheet](/docs/compute/guide/cheat-sheet). ## Choosing a shape | If your workload is | Take | Because | | --- | --- | --- | | Ordinary — a web app, an API, a worker | `cd-standard-*` | 1:2 or 1:4 cores to memory | | CPU-bound — building, encoding, agents | `cd-compute-*` | Twice the cores per gigabyte | | Memory-hungry — a cache, a resident dataset | `cd-memory-*` | 8 GB per core | | Disk-bound | Nothing here | The storage is spinning disk; buy RAM instead | **Memory is the real constraint**, on the box and in your quota: CPU is overcommitted at most 2×, memory is never overcommitted at all. A new account starts at 4 vCPU and 8 GB — [Quotas and limits](/docs/account/cloud/quotas-and-limits). An AWS instance-type name works anywhere ours does: `m5.large` resolves to `cd-standard-2-8`, exactly 2 vCPU and 8 GB. The response names ours, so write ours in Terraform to avoid a perpetual diff. ## When a launch fails | Error | Means | Do | | --- | --- | --- | | `InsufficientInstanceCapacity` | The box cannot fit that shape now | Take a smaller shape, or wait. A quota increase will not help | | `InstanceLimitExceeded` | Your quota | Ask for more | | `InvalidParameterValue` on `InstanceType` | Unknown shape or alias | Check [Instance types](/docs/compute/ec2/ec2-api-instance-types) | | `SignatureDoesNotMatch` | Almost always the region | It is `hel1`, and it must match your credential scope | `GET /v1/catalog` says whether a shape is in stock **right now**, computed from live capacity — [Catalog and availability](/docs/billing/pricing/catalog-and-availability). ## When it boots but does not work The machine reports `running` as soon as it boots. That says nothing about whether your configuration applied: **cloud-init's `runcmd` swallows failures**, so every command in it can fail and the launch still looks clean. Write setup scripts that can be re-run, end them with a sentinel you can check from outside, and when in doubt open the [VNC console](/docs/platform/console/web-console-and-vnc) — it works before the network does. ## Go further - [Developer guide](/docs/compute/guide/developer-guide) — requests, signing, errors - [API reference](/docs/compute/ec2/api-reference-welcome) — every call - [Cheat sheet](/docs/compute/guide/cheat-sheet) — the commands, one page - [Instance types](/docs/compute/ec2/ec2-api-instance-types) - [Machine images](/docs/compute/ec2/ec2-api-machine-images) - [Quotas and limits](/docs/account/cloud/quotas-and-limits) --- # Identity and access — IAM ## Welcome Source: https://shelfcs.com/docs/identity/iam/api-reference-welcome Endpoint, signing region, API version, and the authorisation model in one paragraph. > [!caution] > > **This service is not deployed.** There is no `iam.shelfcs.com` or > `sts.shelfcs.com` in the [published endpoint > list](/docs/platform/endpoints/service-endpoints), and no call on this page > will answer. This page is the specification, not a description of something > running. > > For the identity system that does exist — the identity service users, one project, one > role, and EC2 access keys — read [Identity, as > deployed](/docs/identity/keystone/identity-as-deployed). ## Objective ## Endpoint **PROPOSED:** IAM is a global service. One endpoint, and a fixed signing region. iam. The signing region for IAM is `hel1` regardless of where the caller is, because a global service still requires a region in the credential scope. **OPEN (Michael):** whether a global IAM is correct, or whether identities are regional. A global IAM means one identity store, which is what customers expect; it also means the identity store is a single failure domain for every region. ## API version **PROPOSED:** `2010-05-08`, the AWS IAM API version, so that existing clients send a version we recognise. ## What IAM does - Holds **users**: identities within an account - Issues **access keys**: the credentials that sign API requests - Holds **roles**: identities that can be assumed, with temporary credentials - Holds **policies**: the documents that decide whether a request is allowed - Answers **authorisation**: for every request to every service ## What IAM does not do - It is not an identity provider for a customer's own application's end users. - It does not federate with an external identity provider in v1. - It does not support cross-account access in v1. - It does not manage sign-in to the console. Console sign-in is an Account concern; API credentials are an IAM concern, and the two are separate. ## Authorisation model in one paragraph A request arrives signed by an access key. The key identifies a user or a role session in exactly one account. Every policy attached to that identity is gathered, and the request's action and resource are evaluated against them. An explicit deny wins over any allow. In the absence of an allow, the request is denied. Nothing is permitted by default, including for the account owner's own resources. ## Go further - Shelf Cloud API conventions - Shelf Cloud API Reference --- ## Access key actions Source: https://shelfcs.com/docs/identity/iam/api-reference-access-keys CreateAccessKey, ListAccessKeys, UpdateAccessKey, DeleteAccessKey, GetAccessKeyLastUsed. > [!caution] > > **This service is not deployed.** There is no `iam.shelfcs.com` or > `sts.shelfcs.com` in the [published endpoint > list](/docs/platform/endpoints/service-endpoints), and no call on this page > will answer. This page is the specification, not a description of something > running. > > For the identity system that does exist — the identity service users, one project, one > role, and EC2 access keys — read [Identity, as > deployed](/docs/identity/keystone/identity-as-deployed). ## Objective Access keys are the credentials that sign API requests. These five actions issue them, list them, disable them, delete them, and tell you whether one is still in use. ## Permissions | Action | Resource scope | | --- | --- | | `iam:CreateAccessKey` | `user/` | | `iam:ListAccessKeys` | `user/` | | `iam:UpdateAccessKey` | `user/` | | `iam:DeleteAccessKey` | `user/` | | `iam:GetAccessKeyLastUsed` | `user/` | Called with no `UserName`, each acts on the calling identity — which means a credential able to call `CreateAccessKey` on itself can mint further credentials for itself. Scope this action deliberately. ## Instructions ### CreateAccessKey | Parameter | Type | Required | Notes | | --- | --- | --- | --- | | `UserName` | string | No | Defaults to the calling identity | **Response** | Element | Notes | | --- | --- | | `AccessKeyId` | Sent in the clear on every request thereafter | | `SecretAccessKey` | **Returned here and never again** | | `UserName` | | | `Status` | `Active` | | `CreateDate` | | ```xml deploy AKIAIOSFODNN7EXAMPLE Active wJalrXUtnFEMI/K7MDENG/bPxRfiCYEXAMPLEKEY 2026-09-03T14:22:07Z b1e2c3d4-5678-90ab-cdef-1234567890ab ``` > [!warning] > > The secret is not stored in a form we can return. Support cannot retrieve it. > If it is lost, delete the key and create another. **PROPOSED:** two keys per user maximum. That limit is not arbitrary — it is exactly what makes zero-downtime rotation possible, and it is why the limit is two rather than one. **Not idempotent.** A retry after a lost response mints a second key, and the first secret is unrecoverable. An automated caller should list keys before creating one. **Errors:** `LimitExceeded` (409) when the user already holds two keys, `NoSuchEntity` (404). ### ListAccessKeys | Parameter | Type | Notes | | --- | --- | --- | | `UserName` | string | Defaults to the caller | | `MaxItems`, `Marker` | | Paginated | Returns `AccessKeyId`, `Status`, `UserName`, `CreateDate` for each key. **Never returns a secret**, because none is stored. ### UpdateAccessKey | Parameter | Type | Required | Notes | | --- | --- | --- | --- | | `AccessKeyId` | string | Yes | | | `Status` | string | Yes | `Active` or `Inactive` | | `UserName` | string | No | | Deactivating is the reversible half of deletion, and it is the reason rotation is safe: - An inactive key fails closed. Anything still using it breaks visibly, with `InvalidClientTokenId`, rather than continuing silently. - It can be reactivated. Deletion cannot be undone. **PROPOSED:** we publish the propagation time for a status change. A customer responding to an exposed key needs to know when the revocation is actually in force, and "eventually" is not an answer at that moment. Requests already in flight and correctly signed will complete. ### DeleteAccessKey | Parameter | Type | Required | | --- | --- | --- | | `AccessKeyId` | string | Yes | | `UserName` | string | No | Permanent. The key id is never reused. Deleting a key that does not exist returns `NoSuchEntity`. ### GetAccessKeyLastUsed | Parameter | Type | Required | | --- | --- | --- | | `AccessKeyId` | string | Yes | **Response:** `UserName`, and a `AccessKeyLastUsed` structure with `LastUsedDate`, `ServiceName` and `Region`. `LastUsedDate` is absent if the key has never been used. This action exists so that a key can be retired on evidence rather than on hope. It is the only way to answer "is anything still using this?" **PROPOSED:** the value is updated asynchronously and may lag. We publish the bound, because a customer deleting a key relies on it: a stale timestamp is only evidence of disuse if the lag is much shorter than the observation window. ### Rotation, in order 1. `CreateAccessKey` — a second key for the same user 2. Deploy it everywhere the old one is used 3. `UpdateAccessKey` with `Status=Inactive` on the old key 4. Wait, and watch for `InvalidClientTokenId` failures 5. `GetAccessKeyLastUsed` on the old key to confirm disuse 6. `DeleteAccessKey` Steps 3 and 5 are the technique. Skipping straight to deletion turns a reversible mistake into an outage. ## Go further - Access keys - [User and group actions](/docs/identity/iam/api-reference-users-and-groups) --- ## Actions Source: https://shelfcs.com/docs/identity/iam/api-reference-actions GetAccessKeyLastUsed exists so that a customer can rotate keys safely: it is > [!caution] > > **This service is not deployed.** There is no `iam.shelfcs.com` or > `sts.shelfcs.com` in the [published endpoint > list](/docs/platform/endpoints/service-endpoints), and no call on this page > will answer. This page is the specification, not a description of something > running. > > For the identity system that does exist — the identity service users, one project, one > role, and EC2 access keys — read [Identity, as > deployed](/docs/identity/keystone/identity-as-deployed). ## Objective ## Users | Action | Purpose | | --- | --- | | `CreateUser` | Create a user in this account | | `GetUser` | Retrieve a user | | `ListUsers` | List users, paginated | | `UpdateUser` | Change a user's name or path | | `DeleteUser` | Delete a user; fails if credentials or attachments remain | ## Groups | Action | Purpose | | --- | --- | | `CreateGroup` | Create a group | | `GetGroup` | Retrieve a group and its members | | `ListGroups` | List groups, paginated | | `DeleteGroup` | Delete a group | | `AddUserToGroup` | Add a member | | `RemoveUserFromGroup` | Remove a member | | `ListGroupsForUser` | Groups a user belongs to | ## Access keys | Action | Purpose | | --- | --- | | `CreateAccessKey` | Issue a key pair; the secret is returned once | | `ListAccessKeys` | List key ids for a user; never returns secrets | | `UpdateAccessKey` | Activate or deactivate a key | | `DeleteAccessKey` | Delete a key permanently | | `GetAccessKeyLastUsed` | When and against which service a key was last used | `GetAccessKeyLastUsed` exists so that a customer can rotate keys safely: it is the only way to know whether a key is still in use before deleting it. ## Roles | Action | Purpose | | --- | --- | | `CreateRole` | Create a role with a trust policy | | `GetRole` | Retrieve a role | | `ListRoles` | List roles, paginated | | `UpdateAssumeRolePolicy` | Replace the trust policy | | `DeleteRole` | Delete a role | **OPEN (Michael):** whether roles and temporary credentials ship in v1 at all. If they do, `AssumeRole` and the credential-vending endpoint are part of v1 and need their own decisions on session duration and token format. If they do not, these actions are removed rather than stubbed. ## Policies | Action | Purpose | | --- | --- | | `CreatePolicy` | Create a managed policy | | `GetPolicy` | Retrieve a managed policy and its metadata | | `GetPolicyVersion` | Retrieve a specific version's document | | `ListPolicies` | List managed policies, paginated | | `CreatePolicyVersion` | Add a version, optionally setting it default | | `DeletePolicy` | Delete a managed policy | | `AttachUserPolicy` | Attach a managed policy to a user | | `DetachUserPolicy` | Detach it | | `AttachGroupPolicy` | Attach to a group | | `DetachGroupPolicy` | Detach it | | `AttachRolePolicy` | Attach to a role | | `DetachRolePolicy` | Detach it | | `ListAttachedUserPolicies` | What is attached to a user | | `PutUserPolicy` | Write an inline policy on a user | | `GetUserPolicy` | Read an inline policy | | `DeleteUserPolicy` | Remove an inline policy | **PROPOSED:** managed policies and inline policies both exist, because customer tooling assumes both. Managed policies are versioned; inline policies are not. ## Account-level | Action | Purpose | | --- | --- | | `GetCallerIdentity` | Who am I, which account, which ARN | `GetCallerIdentity` is the first call anyone makes. It must work with any valid credential, require no permission, and never fail for authorisation reasons — it is how a customer proves their credentials are configured correctly. **PROPOSED:** `GetCallerIdentity` lives at the STS endpoint in AWS, not IAM. Whether we mirror that split, or answer it from IAM, is a compatibility decision: clients call `sts.` for it. Mirroring AWS means shipping an STS endpoint in v1 even if nothing else in STS exists. ## Not in v1 Federation, identity providers, SAML, OIDC, MFA devices, service-linked roles, account aliases, credential reports, permission boundaries, and organisations. They are named here so that their absence is deliberate rather than an oversight, and so that nothing in v1 makes them harder to add. ## Go further - Shelf Cloud API conventions - Shelf Cloud API Reference --- ## Data types Source: https://shelfcs.com/docs/identity/iam/api-reference-data-types Every type the IAM API returns, field by field, with mutability stated. > [!caution] > > **This service is not deployed.** There is no `iam.shelfcs.com` or > `sts.shelfcs.com` in the [published endpoint > list](/docs/platform/endpoints/service-endpoints), and no call on this page > will answer. This page is the specification, not a description of something > running. > > For the identity system that does exist — the identity service users, one project, one > role, and EC2 access keys — read [Identity, as > deployed](/docs/identity/keystone/identity-as-deployed). ## Objective ## User | Field | Type | Notes | | --- | --- | --- | | `UserName` | string | Required. Mutable. Unique within the account. | | `UserId` | string | Immutable. Opaque. | | `Arn` | string | Immutable. `arn:aws-shelf:iam:::user/` | | `Path` | string | Optional. Mutable. Defaults to `/`. | | `CreateDate` | timestamp | Immutable. | **PROPOSED:** `UserName` constraints — 1 to 64 characters, alphanumeric plus `+=,.@_-`. This matches what existing tooling validates before it calls us, so a narrower rule would reject names that clients believe are valid. Note that the ARN embeds the user name, which is mutable. Renaming a user therefore changes its ARN, and any policy referencing the old ARN stops matching. This is inherited AWS behaviour, it is a genuine hazard, and it is documented rather than silently fixed, because fixing it would break compatibility. ## Group | Field | Type | Notes | | --- | --- | --- | | `GroupName` | string | Required. Mutable. Unique within the account. | | `GroupId` | string | Immutable. | | `Arn` | string | Immutable form; embeds the name. | | `Path` | string | Optional. Mutable. | | `CreateDate` | timestamp | Immutable. | ## AccessKey | Field | Type | Notes | | --- | --- | --- | | `AccessKeyId` | string | Immutable. | | `SecretAccessKey` | string | Returned only by `CreateAccessKey`, once. | | `Status` | string | `Active` or `Inactive`. Mutable. | | `UserName` | string | Immutable. The owning user. | | `CreateDate` | timestamp | Immutable. | A secret is never stored in a form we can return. `ListAccessKeys` returns ids and status only. ## AccessKeyLastUsed | Field | Type | Notes | | --- | --- | --- | | `LastUsedDate` | timestamp | Absent if never used. | | `ServiceName` | string | The service last called. | | `Region` | string | The region of that call. | **PROPOSED:** last-used is updated asynchronously and may lag. The lag is documented with a bound, because a customer deleting a key relies on it. ## Role | Field | Type | Notes | | --- | --- | --- | | `RoleName` | string | Required. Mutable. | | `RoleId` | string | Immutable. | | `Arn` | string | Immutable form; embeds the name. | | `AssumeRolePolicyDocument` | string | Required. Mutable. | | `MaxSessionDuration` | integer | Seconds. **OPEN**: bounds. | | `CreateDate` | timestamp | Immutable. | ## Policy | Field | Type | Notes | | --- | --- | --- | | `PolicyName` | string | Required. Immutable after creation. | | `PolicyId` | string | Immutable. | | `Arn` | string | Immutable. | | `DefaultVersionId` | string | Mutable. | | `AttachmentCount` | integer | Derived. | | `CreateDate`, `UpdateDate` | timestamp | | **PROPOSED:** a managed policy keeps up to five versions, matching AWS, so that tooling which rotates versions does not fail on the sixth. ## Go further - Shelf Cloud API conventions - Shelf Cloud API Reference --- ## IAM API errors Source: https://shelfcs.com/docs/identity/iam/api-reference-errors The IAM error envelope and every code the service returns. > [!caution] > > **This service is not deployed.** There is no `iam.shelfcs.com` or > `sts.shelfcs.com` in the [published endpoint > list](/docs/platform/endpoints/service-endpoints), and no call on this page > will answer. This page is the specification, not a description of something > running. > > For the identity system that does exist — the identity service users, one project, one > role, and EC2 access keys — read [Identity, as > deployed](/docs/identity/keystone/identity-as-deployed). ## Objective **Read this page to handle IAM errors correctly.** IAM uses the standard query error envelope, which differs from EC2's. See [IAM requests and responses](/docs/identity/iam/api-reference-making-requests) for the shape. ## Requirements - None ## Instructions ### Authentication and authorisation | Code | Status | `Type` | Cause | | --- | --- | --- | --- | | `IncompleteSignature` | 400 | Sender | The `Authorization` header is malformed | | `SignatureDoesNotMatch` | 403 | Sender | The signature does not match the request | | `InvalidClientTokenId` | 403 | Sender | The access key id is unknown or inactive | | `MissingAuthenticationToken` | 403 | Sender | The request was not signed | | `RequestExpired` | 400 | Sender | The timestamp is outside the permitted skew | | `AccessDenied` | 403 | Sender | Policy does not allow this action | `AccessDenied` is the outcome clients branch on most often, which is why it is first in this table rather than buried below the service-specific codes. ### Service-specific | Code | Status | `Type` | Cause | | --- | --- | --- | --- | | `NoSuchEntity` | 404 | Sender | The user, group, role, policy or key does not exist | | `EntityAlreadyExists` | 409 | Sender | A resource with that name already exists | | `DeleteConflict` | 409 | Sender | The resource still has attachments or credentials | | `LimitExceeded` | 409 | Sender | An account quota would be exceeded | | `MalformedPolicyDocument` | 400 | Sender | The policy document is not valid | | `InvalidInput` | 400 | Sender | A parameter value is not acceptable | | `UnmodifiableEntity` | 409 | Sender | The resource cannot be changed | ### Throttling and server errors | Code | Status | `Type` | Cause | Retry | | --- | --- | --- | --- | --- | | `Throttling` | 400 | Sender | Request rate exceeded | Yes, with backoff | | `ServiceFailure` | 500 | Receiver | An unexpected failure on our side | Yes | `Throttling` is a 400 here and `RequestLimitExceeded` is a 503 in EC2. That inconsistency is inherited: each service matches the code and status its AWS counterpart returns, because client retry logic keys on both. `ServiceFailure` is the code AWS IAM documents on every action for an unexpected error, so it is the code used here rather than a name of our own. ### Policy evaluation failure A policy that cannot be evaluated is our failure, not a denial. It returns `ServiceFailure` with a 500 — never `AccessDenied`. The distinction matters more than it appears: reported as a denial, a customer spends a day rewriting a policy that was correct all along. ### Absence versus denial An entity in another account, or one that has never existed, returns `NoSuchEntity`. An entity in the caller's own account that policy forbids returns `AccessDenied`. Existence is never confirmed to a caller not entitled to know it. ### Deletion conflicts `DeleteConflict` rather than a cascading delete is deliberate. Deleting a user must not silently delete the access keys something is still authenticating with: the caller is told what remains and deletes it explicitly. ## Go further - [IAM requests and responses](/docs/identity/iam/api-reference-making-requests) - [Actions](/docs/identity/iam/api-reference-actions) - IAM troubleshooting --- ## IAM requests, responses and pagination Source: https://shelfcs.com/docs/identity/iam/api-reference-making-requests The IAM wire format, its response envelopes, and its pagination model. > [!caution] > > **This service is not deployed.** There is no `iam.shelfcs.com` or > `sts.shelfcs.com` in the [published endpoint > list](/docs/platform/endpoints/service-endpoints), and no call on this page > will answer. This page is the specification, not a description of something > running. > > For the identity system that does exist — the identity service users, one project, one > role, and EC2 access keys — read [Identity, as > deployed](/docs/identity/keystone/identity-as-deployed). ## Objective IAM speaks the standard AWS query protocol. It is **not** the same protocol EC2 speaks, and the differences are not cosmetic: different response envelope, different request-id element, different pagination parameters. **Read this page before implementing an IAM client or the IAM service.** ## Requirements - An access key id and secret access key ## Instructions ### Request structure Form-encoded `POST`, with the action and version as parameters: ``` POST / HTTP/1.1 Host: iam.shelfcloud.com Content-Type: application/x-www-form-urlencoded; charset=utf-8 X-Amz-Date: 20260903T142207Z Authorization: AWS4-HMAC-SHA256 Credential=AKIA.../20260903/hel1/iam/aws4_request, ... Action=ListUsers&Version=2010-05-08&MaxItems=100 ``` ### Success envelope ```xml true deploy AIDA0A1B2C3D4E5F60718 arn:aws-shelf:iam::123456789012:user/deploy / 2026-09-03T14:22:07Z eyJwYWdlIjoyfQ b1e2c3d4-5678-90ab-cdef-1234567890ab ``` Three things a client depends on and an implementer must get exactly right: - The result is wrapped in `` inside ``. - List members are `` elements, not elements named after the type. - The request id lives in `` — **`RequestId`**, with a lowercase `d`, unlike EC2's ``. ### Error envelope ```xml Sender NoSuchEntity The user with name deploy cannot be found. b1e2c3d4-5678-90ab-cdef-1234567890ab ``` `` is `Sender` for client errors and `Receiver` for server errors. EC2's envelope has no ``; IAM's does. A client written for one cannot parse the other. ### Pagination: Marker, not NextToken **IAM does not use `MaxResults` and `NextToken`.** It uses: | Parameter | Direction | Meaning | | --- | --- | --- | | `MaxItems` | Request | Items per page. Default 100, maximum 1000 | | `Marker` | Request | The marker from the previous response | | `IsTruncated` | Response | `true` when more items remain | | `Marker` | Response | Present when `IsTruncated` is `true` | | `PathPrefix` | Request | Restrict to a path prefix | This is not a choice. Every AWS SDK's IAM paginators key on `Marker`/`IsTruncated`. An implementation that returned `NextToken` here would have its token discarded by the client, which would then report only the first page — a silently truncated list of users, in the service that decides who can do what. ```python marker = None while True: kwargs = {"MaxItems": 1000} if marker: kwargs["Marker"] = marker page = iam.list_users(**kwargs) handle(page["Users"]) if not page.get("IsTruncated"): break marker = page["Marker"] ``` The EC2 pagination rules do not apply to IAM. They are documented separately and each page says which service it governs. ### Signing The signing name is `iam`. **PROPOSED:** the signing region is `hel1`, **and `us-east-1` is also accepted.** Both are accepted because IAM is a global service in AWS with `us-east-1` hardcoded in the SDK endpoint rulesets. When an endpoint is overridden, some SDK versions still sign with `us-east-1` and others sign with the caller's configured region. Accepting only one produces signature failures that vary by SDK and by version, which is worse than either answer alone. ### Consistency IAM is eventually consistent. A credential or policy change may take a short time to be visible to every service. **PROPOSED:** we publish the propagation bound, because a customer revoking a compromised key needs to know when the revocation is actually in force. ## Go further - [Errors](/docs/identity/iam/api-reference-errors) - [Actions](/docs/identity/iam/api-reference-actions) - Signing requests --- ## Policy actions Source: https://shelfcs.com/docs/identity/iam/api-reference-policies Managed policies, inline policies, attachment, and policy versions. > [!caution] > > **This service is not deployed.** There is no `iam.shelfcs.com` or > `sts.shelfcs.com` in the [published endpoint > list](/docs/platform/endpoints/service-endpoints), and no call on this page > will answer. This page is the specification, not a description of something > running. > > For the identity system that does exist — the identity service users, one project, one > role, and EC2 access keys — read [Identity, as > deployed](/docs/identity/keystone/identity-as-deployed). ## Objective Policies decide what an identity may do. These actions create them, version them, attach them and read them back. ## Permissions | Action | Resource scope | | --- | --- | | `iam:CreatePolicy`, `iam:DeletePolicy` | `policy/` | | `iam:GetPolicy`, `iam:GetPolicyVersion`, `iam:ListPolicies` | `policy/*` | | `iam:CreatePolicyVersion`, `iam:SetDefaultPolicyVersion` | `policy/` | | `iam:AttachUserPolicy`, `iam:DetachUserPolicy` | Both `user/` and `policy/` | | `iam:PutUserPolicy`, `iam:DeleteUserPolicy` | `user/` | The attach and detach actions evaluate against both the identity and the policy. > [!warning] > > An identity holding `iam:AttachUserPolicy` and `iam:CreatePolicy` can grant > itself any permission in the account. There is no permission boundary in v1 > to constrain that. Treat those two actions as equivalent to full > administrative access, and grant them accordingly. ## Instructions ### Managed and inline **Managed policies** are standalone documents with their own ARN, attachable to many identities, and versioned. Use them for anything reused. **Inline policies** are embedded in one identity, have no ARN of their own, are not versioned, and are deleted with the identity. Use them for permissions that belong to exactly one identity and should not outlive it. ### CreatePolicy | Parameter | Type | Required | Notes | | --- | --- | --- | --- | | `PolicyName` | string | Yes | Immutable after creation | | `PolicyDocument` | string | Yes | URL-encoded JSON | | `Path` | string | No | | | `Description` | string | No | Immutable after creation | **PROPOSED:** a policy document is at most 6144 characters, counted with whitespace removed. The document is validated at creation. A document that does not parse, names a malformed ARN, or uses a construct we do not implement is rejected with `MalformedPolicyDocument` — never accepted and then evaluated differently from what it says. That rule matters most for any construct in the policy language we do not support: silently ignoring an unsupported `Condition` would turn a restrictive policy into a permissive one, which is the worst possible failure mode for this service. **Response:** the `Policy` — `PolicyName`, `PolicyId`, `Arn`, `DefaultVersionId`, `AttachmentCount`, `CreateDate`, `UpdateDate`. ### Versions A managed policy is not edited in place. `CreatePolicyVersion` adds a version, optionally making it the default: | Parameter | Type | Required | Notes | | --- | --- | --- | --- | | `PolicyArn` | string | Yes | | | `PolicyDocument` | string | Yes | | | `SetAsDefault` | boolean | No | Default `false` | **PROPOSED:** five versions per policy, matching what tooling expects. The sixth returns `LimitExceeded`; delete an old version first. Only the **default** version is evaluated. A new version created without `SetAsDefault` changes nothing until `SetDefaultPolicyVersion` is called — which is what makes a policy change reviewable before it takes effect, and instantly reversible afterwards by setting the previous version back. ### Attachment | Action | Attaches to | | --- | --- | | `AttachUserPolicy` / `DetachUserPolicy` | A user | | `AttachGroupPolicy` / `DetachGroupPolicy` | A group | | `AttachRolePolicy` / `DetachRolePolicy` | A role | | `ListAttachedUserPolicies` | Lists what is attached to a user | Attaching an already-attached policy succeeds and changes nothing. A policy cannot be deleted while attached to anything: `DeletePolicy` returns `DeleteConflict` until every attachment is removed. ### Inline policies | Action | Purpose | | --- | --- | | `PutUserPolicy` | Write an inline policy; overwrites by name | | `GetUserPolicy` | Read one | | `ListUserPolicies` | Names of inline policies on a user | | `DeleteUserPolicy` | Remove one | `PutUserPolicy` overwrites without warning. There is no version history and no undo — the previous document is gone. This is the practical reason to prefer managed policies for anything that matters. ### Reading a policy back `GetPolicy` returns metadata only, not the document. `GetPolicyVersion` returns the document, URL-encoded. To answer "what can this user do", a caller must gather: inline policies on the user, managed policies attached to the user, and both again for every group the user belongs to. There is no single call that returns an effective permission set, and no policy simulator in v1. **OPEN (Michael):** whether a policy simulator exists in v1. Without one, the only way to test a policy is to hold the credential and try the call — which is what `DryRun` on the EC2 actions is for. ## Go further - [Policy reference](/docs/identity/iam/api-reference-policy-reference) - Policies - [IAM errors](/docs/identity/iam/api-reference-errors) --- ## Policy reference Source: https://shelfcs.com/docs/identity/iam/api-reference-policy-reference This page is the most consequential in the whole v1 documentation set. Every > [!caution] > > **This service is not deployed.** There is no `iam.shelfcs.com` or > `sts.shelfcs.com` in the [published endpoint > list](/docs/platform/endpoints/service-endpoints), and no call on this page > will answer. This page is the specification, not a description of something > running. > > For the identity system that does exist — the identity service users, one project, one > role, and EC2 access keys — read [Identity, as > deployed](/docs/identity/keystone/identity-as-deployed). ## Objective This page is the most consequential in the whole v1 documentation set. Every service's actions have to be nameable in a policy on the day that service ships, so the naming scheme is decided once, here, and never changed. ## Document structure **PROPOSED:** the AWS IAM policy JSON document structure, unchanged, so that existing policies and existing policy tooling work. { "Version": "2012-10-17", "Statement": [ { "Sid": "AllowReadOnlyEC2", "Effect": "Allow", "Action": ["ec2:DescribeInstances", "ec2:DescribeImages"], "Resource": "*" } ] } ## Action naming An action is `:`, where the service prefix is the same string used in the endpoint host and in the SigV4 credential scope. One string, three places, always identical. **PROPOSED:** v1 service prefixes — `iam`, `ec2`, `account`, `billing`. Prefixes for services that do not exist yet are reserved now: `s3`, `elb`, `rds`, `ecr`, `sqs`, `logs`, `route53`, `acm`, `secretsmanager`, `kms`. They are reserved so that a future service cannot be forced into a name that clashes with an existing one, and so that AWS-compatible policies keep working when the matching service arrives. Wildcards are permitted in actions: `ec2:Describe*`, `ec2:*`, `*`. ## Resource matching Resources are matched by ARN, with `*` and `?` wildcards permitted in any segment. arn:aws-shelf:ec2:hel1:123456789012:instance/* arn:aws-shelf:ec2:*:123456789012:instance/i-0a1b2c3d4e5f60718 An action that operates on no particular resource — a list across the account, for example — is documented per action as requiring `Resource: "*"`, and the per-action page states it explicitly. Guessing which actions are resource-scoped is the single largest source of policy mistakes, so every action page carries the answer. ## Evaluation 1. An explicit `Deny` in any applicable policy denies the request. 2. Otherwise, an `Allow` in any applicable policy allows it. 3. Otherwise, the request is denied. There is no implicit allow, for anyone, including the account owner. **OPEN (Michael):** whether the account owner, or a root credential, bypasses policy evaluation. AWS's root user does. A design where nothing bypasses policy is safer and is also a way to lock an account out of itself permanently. Both positions are defensible; the decision must be made before the first policy is written. ## Conditions **OPEN (Michael):** whether condition keys exist in v1, and which. Conditions are how customers express "only from this address" and "only with MFA". They are also a large surface. Shipping a policy engine without conditions and adding them later is safe; shipping them incompletely is not. ## Implementation **PROPOSED:** the policy engine is Cedar rather than hand-rolled. Cedar is an open-source policy language with a Go implementation, designed for exactly this evaluation model, and it removes the risk of a hand-written evaluator that subtly disagrees with itself between services. If Cedar is chosen, the AWS-shaped policy document is translated into Cedar at write time, and the translation is part of the contract: a policy that cannot be translated must be rejected at creation, never silently accepted and evaluated differently. ## Go further - Shelf Cloud API conventions - Shelf Cloud API Reference --- ## User and group actions Source: https://shelfcs.com/docs/identity/iam/api-reference-users-and-groups CreateUser, GetUser, ListUsers, UpdateUser, DeleteUser, and the group actions. > [!caution] > > **This service is not deployed.** There is no `iam.shelfcs.com` or > `sts.shelfcs.com` in the [published endpoint > list](/docs/platform/endpoints/service-endpoints), and no call on this page > will answer. This page is the specification, not a description of something > running. > > For the identity system that does exist — the identity service users, one project, one > role, and EC2 access keys — read [Identity, as > deployed](/docs/identity/keystone/identity-as-deployed). ## Objective A user is an identity within one account — a person, or a machine. A group is a collection of users that permissions are attached to once rather than repeatedly. ## Permissions | Action | Resource scope | | --- | --- | | `iam:CreateUser` | `user/` | | `iam:GetUser`, `iam:ListUsers` | `user/*` | | `iam:UpdateUser`, `iam:DeleteUser` | `user/` | | `iam:CreateGroup`, `iam:DeleteGroup` | `group/` | | `iam:AddUserToGroup`, `iam:RemoveUserFromGroup` | Both `user/` and `group/` | `AddUserToGroup` evaluates against both resources. A policy naming only the group denies the call. ## Instructions ### CreateUser | Parameter | Type | Required | Notes | | --- | --- | --- | --- | | `UserName` | string | Yes | Unique within the account | | `Path` | string | No | Defaults to `/` | | `Tags.member.N` | list | No | | **PROPOSED:** `UserName` is 1 to 64 characters, alphanumeric plus `+=,.@_-`. This matches what client tooling validates before calling, so a narrower rule would reject names the client believes are valid. **Response:** the `User` — `UserName`, `UserId`, `Arn`, `Path`, `CreateDate`. A new user can do nothing. It has no credentials and no policies, and every call it could make is denied until both exist. That is the intended starting state. **Errors:** `EntityAlreadyExists` (409), `InvalidInput` (400), `LimitExceeded` (409). ### GetUser | Parameter | Type | Required | Notes | | --- | --- | --- | --- | | `UserName` | string | No | Defaults to the calling identity | Called with no argument it returns the caller. That is the cheapest way for a credential to discover what it is, and it is what tooling does at startup. ### ListUsers | Parameter | Type | Notes | | --- | --- | --- | | `PathPrefix` | string | Restrict to a path | | `MaxItems` | integer | Default 100, maximum 1000 | | `Marker` | string | From a previous response | **Paginated with `Marker` and `IsTruncated`, not `NextToken`.** See [IAM requests and responses](/docs/identity/iam/api-reference-making-requests). Building the EC2 pagination shape here truncates the user list silently. ### UpdateUser | Parameter | Type | Required | | --- | --- | --- | | `UserName` | string | Yes | | `NewUserName` | string | No | | `NewPath` | string | No | > [!warning] > > **Renaming a user changes its ARN**, because the ARN embeds the name. Every > policy that references the old ARN stops matching, silently — the request is > simply denied as if no policy granted it. > > This is inherited behaviour and we keep it for compatibility. Before renaming, > find every policy referencing the old ARN. Attaching policies to groups rather > than to users avoids the problem entirely. ### DeleteUser | Parameter | Type | Required | | --- | --- | --- | | `UserName` | string | Yes | Fails with `DeleteConflict` while the user still has access keys, attached policies, inline policies, or group memberships. This is deliberate rather than an inconvenience. A cascading delete would silently destroy the credential something is still authenticating with, and the first anyone would know is a production outage. The caller is told what remains and removes it explicitly. The order that works: 1. `DeleteAccessKey` for every key — check `GetAccessKeyLastUsed` first 2. `DetachUserPolicy` for every managed policy 3. `DeleteUserPolicy` for every inline policy 4. `RemoveUserFromGroup` for every group 5. `DeleteUser` ### Group actions | Action | Purpose | | --- | --- | | `CreateGroup` | Create a group | | `GetGroup` | The group and its members, paginated by `Marker` | | `ListGroups` | Groups in the account | | `ListGroupsForUser` | Groups one user belongs to | | `AddUserToGroup` | Add a member | | `RemoveUserFromGroup` | Remove a member | | `DeleteGroup` | Delete; fails with `DeleteConflict` while members remain | Groups do not nest. A group cannot contain another group. **Attach policies to groups rather than to users.** A user's permissions then change by moving them between groups, which is auditable, reversible, and does not require editing a policy document under time pressure. ## Go further - [Access key actions](/docs/identity/iam/api-reference-access-keys) - [Policy actions](/docs/identity/iam/api-reference-policies) - Policies --- # Identity and access — KEYSTONE ## Identity, as deployed Source: https://shelfcs.com/docs/identity/keystone/identity-as-deployed What actually authenticates you today — the identity service users, projects and one role — and how each AWS IAM concept maps onto it, or does not. ## Objective **The IAM and STS sections of this reference describe a service that is not deployed.** There is no `iam.shelfcs.com` in the [published endpoint list](/docs/platform/endpoints/service-endpoints), and no call in those pages will answer. This page describes the identity system that does exist. Where it disagrees with an IAM or STS page, this one is what will answer your call. `an internal service`, `an internal service`. ## Two identity systems, one account | | Storefront identity | Cloud identity | | --- | --- | --- | | Who runs it | the hosted auth service (Better Auth) | the identity service | | You are | Your email address | `cust-` | | Credential | Password, then a 15-minute JWT | Cloud password, and EC2 access keys | | Gets you | Signup, billing, catalog, the Account API | Every cloud API and both consoles | | Created | When you sign up | At `POST /v1/onboard` | They are two accounts created together, and they do **not** share a session. Neither is IAM. ## What the identity service identity is made of | Object | What it is | How many you get | | --- | --- | --- | | **Project** | The tenancy boundary. Every instance, volume, image and key belongs to one | Exactly one, named `cust-` | | **User** | The login. Holds the password and the EC2 credentials | Exactly one | | **Role** | What the user may do on the project | Exactly one: `member` | | **Domain** | the identity service's namespace above projects | `Default`, always | | **EC2 credential** | The access key pair SigV4 signs with | As many as you mint — [Access keys](/docs/account/cloud/access-keys) | That is the whole model. One user, one project, one role. ## Every IAM concept, and what it maps to | AWS IAM | Here | | --- | --- | | AWS account | the identity service **project** | | IAM user | The **one** the identity service user. You cannot create a second | | IAM group | **Does not exist** | | IAM role, `AssumeRole` | **Does not exist.** There is no STS, no role assumption, no cross-account trust | | Managed or inline policy | **Does not exist.** Authorisation is the `member` role and the platform's own policy files, which you cannot edit | | Policy condition keys, `aws:SourceIp`, MFA conditions | **Do not exist** | | Permission boundary, SCP | **Do not exist** | | Instance profile | **Does not exist.** An instance carries no identity of its own and cannot call the API as itself | | Access key | the identity service **EC2 credential** — the one concept that survives intact | | Temporary credentials | **Do not exist.** Every credential is long-lived until deleted | | `iam:PassRole` | Not applicable | | Account root user | The operator, not you | ## What this means in practice Read these as consequences, not as apology: - **You cannot give a teammate narrower access than your own.** There is one user. Sharing access means sharing that user's password or one of its keys. - **You cannot give CI a read-only credential.** A key inherits the whole `member` role — see [Access keys](/docs/account/cloud/access-keys). - **You cannot separate staging from production inside one account.** The project is the only boundary, and you get one. Separation means **two accounts**, signed up separately. - **An application on an instance cannot authenticate as the instance.** There is no metadata-served role. If it needs to call the API, it needs a key on disk, with everything that implies. - **There is no MFA**, on either identity system. - **There is no audit log you can read.** the identity service and the services log server-side; nothing surfaces it to you. Policies written for AWS do not restrict anything here. If your security model depends on IAM policy, this platform does not implement it — and it is better you know that from this page than infer it from a policy that silently never applied. ## What the `member` role permits Everything within your own project: launch, describe and terminate instances; create, attach and delete volumes and snapshots; read images; manage security groups and key pairs; read your own rated usage. Nothing outside it: no creating projects, users, roles, flavors or images, no reading another project, no operator action of any kind. ## The IAM and STS pages They remain published because they are the **specification** — the shape the service will take if and when it is built, written against the AWS model and reviewed. Every one of them now carries a notice saying it is not deployed. Treat them as a design document. Do not write code against them. ## Go further - [Access keys](/docs/account/cloud/access-keys) — the credential that does exist - [Account developer guide](/docs/account/guide/developer-guide) - [Accounts and tenancy](/docs/platform/conventions/accounts-and-tenancy) - [Quotas and limits](/docs/account/cloud/quotas-and-limits) - [Service endpoints](/docs/platform/endpoints/service-endpoints) --- # Identity and access — STS ## Shelf Cloud STS API Reference Source: https://shelfcs.com/docs/identity/sts/api-reference-welcome The identity endpoint every client calls first, and why it exists separately from IAM. > [!caution] > > **This service is not deployed.** There is no `iam.shelfcs.com` or > `sts.shelfcs.com` in the [published endpoint > list](/docs/platform/endpoints/service-endpoints), and no call on this page > will answer. This page is the specification, not a description of something > running. > > For the identity system that does exist — the identity service users, one project, one > role, and EC2 access keys — read [Identity, as > deployed](/docs/identity/keystone/identity-as-deployed). ## Objective STS answers one question in v1: **who am I?** It exists as its own service rather than as an IAM action because that is where every AWS SDK looks for it. `GetCallerIdentity` is modelled on the STS client, signed with the service name `sts`, and sent to an STS endpoint. A client will not find it on IAM no matter where the answer actually lives. **Read this page if you are implementing the service, or configuring a tool that validates credentials at start-up.** ## Requirements - Any valid access key ## Instructions ### Endpoint ``` https://sts.shelfcloud.com ``` **PROPOSED:** STS is global, signed with region `hel1`, and `us-east-1` is also accepted — for the same reason IAM accepts both. See [IAM requests and responses](/docs/identity/iam/api-reference-making-requests). ### API version ``` Version=2011-06-15 ``` ### Protocol The standard query protocol, with the same envelopes as IAM. ### GetCallerIdentity Returns the identity behind the credentials that signed the request. **Request:** no parameters beyond `Action` and `Version`. **Response** | Element | Notes | | --- | --- | | `UserId` | The unique id of the identity | | `Account` | The twelve-digit account id | | `Arn` | The ARN of the identity | ```xml AIDA0A1B2C3D4E5F60718 123456789012 arn:aws-shelf:iam::123456789012:user/deploy b1e2c3d4-5678-90ab-cdef-1234567890ab ``` ### It requires no permissions `GetCallerIdentity` succeeds for any valid credential and cannot be denied by policy. This is deliberate and matters in three places: 1. It is the first call a customer makes, before any policy exists. 2. It is the first diagnostic step in every troubleshooting guide: a failure here means credentials or configuration, not permissions. 3. The Terraform AWS provider calls it during provider configuration unless `skip_credentials_validation` and `skip_requesting_account_id` are both set. Without a working STS endpoint, Terraform fails before it reaches a single resource. ### Not in v1 `AssumeRole`, `GetSessionToken`, `GetFederationToken`, and every other STS action. They return `InvalidAction`. **OPEN (Michael):** whether `AssumeRole` and temporary credentials ship in v1. The IAM role actions depend on this same decision. ### Errors The IAM error envelope and the shared authentication codes. `GetCallerIdentity` returns no service-specific errors, because there is nothing for it to fail at beyond authentication. ## Go further - [Shelf Cloud IAM API Reference](/docs/identity/iam/api-reference-welcome) - Signing requests --- # Platform — CONSOLE ## Web console and VNC Source: https://shelfcs.com/docs/platform/console/web-console-and-vnc the web console at console.shelfcs.com and the noVNC instance console — the two ways to see a machine without SSH. ## Objective Everything on this platform is an API, and an API is no help when an instance will not boot far enough to accept an SSH connection. Two hosts exist for that case. ## The web console ``` https://console.shelfcs.com ``` The platform's web dashboard. Sign in with the cloud username and password issued at onboarding, domain `Default`, and you are in your own project only. What it is good for: - Watching an instance boot, from its console output. - Attaching and detaching volumes when you want to see what you are doing. - Reading the actual error behind a failed launch, which is often more specific than the EC2-shaped code the API is obliged to return. - Finding your project's real quota and usage. What it is not: it is not the way to run infrastructure. Console clicks are unfinished work — anything you want to keep belongs in Terraform against [the EC2 API](/docs/compute/ec2/api-reference-welcome) or the [Account developer guide](/docs/account/guide/developer-guide). `console-api.shelfcs.com` is the web console's own backend. The browser calls it; you do not. ## The instance console ``` https://vnc.shelfcs.com ``` noVNC, reached through a **single-use, time-limited URL** that the compute service mints for one instance: ``` cloud console url show --novnc ``` The URL contains a token, expires quickly, and is not reusable. Do not paste one into a ticket. This is a graphical console attached to the virtual machine's display. It works before the network does, which is the point: a machine with a broken `/etc/fstab`, a firewall rule that locked you out, or a cloud-init failure is reachable here and nowhere else. > [!warning] > > `nofail` in `/etc/fstab` is what keeps you out of this console in the first > place. A volume that is not present at boot, mounted without `nofail`, drops > the machine into emergency mode — and emergency mode wants a root password > that a cloud image does not have. Recovering from that means detaching the > root disk and mounting it on another instance. ## What the EC2 API does not give you | Want | EC2 here | Instead | | --- | --- | --- | | Console text output | `GetConsoleOutput` — **not implemented** | the web console, or `cloud console log show` | | Console screenshot | `GetConsoleScreenshot` — **not implemented** | The VNC console | | Serial console | `SendSerialConsoleSSHPublicKey` — **not implemented** | The VNC console | An AWS-shaped tool that reaches for console output receives `InvalidAction`. This is a real gap in the EC2 layer, not a policy. ## Access, honestly - Both consoles are behind the same ingress as every other endpoint, with the same absence of rate limiting noted on the [Service endpoints](/docs/platform/endpoints/service-endpoints) page. - There is **no MFA** on the cloud login. the identity service's password is the only factor. - There is **no SSO**. Identity for the *storefront* is the hosted auth service; identity for the *cloud* is the identity service; they are two accounts created together at onboarding and they do not share a session. - There is no password-reset endpoint for the cloud password yet. A later `/v1/onboard` call with a different password does **not** change it. ## Go further - [Service endpoints](/docs/platform/endpoints/service-endpoints) - [Account developer guide](/docs/account/guide/developer-guide) - [Account API](/docs/account/cloud/storefront-api-reference) --- # Platform — API conventions ## Accounts and tenancy Source: https://shelfcs.com/docs/platform/conventions/accounts-and-tenancy The account boundary, visibility rules, and the account lifecycle ## Objective **Read this page to understand the single boundary every other rule is drawn around.** ## Requirements - None ## Instructions ### The account is the boundary An account is the unit of tenancy, ownership, authorisation and billing. Every resource belongs to exactly one account, every invoice covers one account, and every credential authenticates into one account. There is no cross-account access in v1. The ARN format reserves room for it without implying it exists. ### Accounts and platform projects **PROPOSED:** an account maps to exactly one project in the underlying platform, one-to-one and permanently. A design in which one account spans several projects, or one project serves several accounts, makes both authorisation and metering ambiguous: a usage record could not be attributed to one payer, and a policy could not be evaluated against one owner. ### Visibility A caller sees only resources in its own account. A resource in another account is not merely forbidden, it is **invisible**: the response is byte-identical to the response for a resource that has never existed. Identifiers cannot be probed by observing which of them return a different error. The complete rule, which every service implements identically: 1. **Another account's resource, or no such resource** → the service's not-found error. The two cases are indistinguishable. 2. **This account, resource exists, policy denies** → `UnauthorizedOperation` or `AccessDenied`. So a denial code confirms existence, and may therefore be returned only within the caller's own account, where existence is not a secret from them. Amazon EC2 behaves this way, returning `InvalidInstanceID.NotFound` for another account's instance, and matching it is both correct and compatible. ### Account ids **PROPOSED:** twelve decimal digits, never beginning with zero. Twelve digits because customer tooling and policy documents assume that shape. A numeric id carries no information about the customer. The leading-zero exclusion is deliberate. An id beginning with zero is parsed as a number by YAML, by spreadsheets, by JSON tooling and by type coercion in infrastructure tools, loses the zero, and then matches nothing. Allocating around the problem costs nothing; discovering it in a customer's pipeline costs a day. Account ids are never reused, including after closure. ### Isolation Resources in different accounts are isolated from one another. **OPEN (Michael):** the isolation guarantee that can honestly be stated in v1. This depends on when per-tenant networking ships. Nothing is claimed until it does — an isolation claim that is not true is the single most damaging sentence this documentation could contain. ### Lifecycle | State | Meaning | API access | Resources | | --- | --- | --- | --- | | `pending` | Created, not verified | Refused | None can exist | | `active` | Normal | Full | Running | | `suspended` | Non-payment or abuse | Refused, except reading the account | Retained, still charged | | `closing` | Closure requested | Refused | Being released | | `closed` | Closed | Refused | Destroyed | A suspended account may still read its own state, so that a blocked customer can find out why they are blocked. Every other action returns `AccountSuspended` with HTTP 403. **OPEN (Michael):** how long a suspended account is retained before closure, and the notice given before anything is destroyed. A customer whose data is about to be deleted is owed a precise answer. **OPEN (Michael):** the retention period for records kept after closure — invoices, tax records and identity details — which is a legal question. ## Go further - [ARNs and identifiers](/docs/platform/conventions/arns-and-identifiers) - [Errors](/docs/platform/conventions/errors) - What is a Shelf Cloud account? --- ## API versioning Source: https://shelfcs.com/docs/platform/conventions/api-versioning How versions are identified per protocol, what is breaking, and how deprecation works ## Objective **Read this page to know what we may change without warning and what we may not.** ## Requirements - None ## Instructions ### How a version is identified It depends on the protocol, and there is no single answer: **Query-protocol services** carry the version as a request parameter: ``` Version=2016-11-15 ``` The SDKs supply it from their own service model; a caller does not set it by hand. **JSON-protocol services carry no version on the wire.** The version is part of the action target: ``` X-Shelf-Target: ShelfBilling_20260101.GetQuote ``` Requiring a `Version` parameter from a JSON-protocol client would break every SDK, because none sends one. ### Which versions are accepted For query-protocol services, **a version at or before the implemented one is accepted.** A later version is rejected with `InvalidParameterValue`. Accepting earlier versions matters: different SDK releases pin different version dates, and hard-rejecting anything but the newest would mean an older but currently supported SDK release fails entirely — contradicting the promise that existing clients work unmodified. For JSON-protocol services, an unknown target is rejected with `InvalidAction`. ### What is a breaking change Never done within a version: - Removing an action, a parameter or a response field - Making an optional parameter required - Narrowing an accepted value range - Changing the meaning of an error code - Changing the type or the unit of a field - Changing the format of an identifier - Changing the HTTP status returned for an existing error code Permitted within a version: - Adding an action - Adding an optional parameter - Adding a response field - Adding a new error code for a genuinely new failure - Widening an accepted range **Clients must ignore response fields they do not recognise.** Stated here explicitly because a client that rejects unknown fields turns every future addition into a breaking change. ### Deprecation **OPEN (Michael):** the deprecation period and the notice given. Whatever is chosen applies to every service, is stated once, and is honoured: an announced date is not brought forward. ### Document history Every document carries a history. Every change to the contract appears there, dated, with what changed. A change that is not written down did not happen. The history is what a customer reads when working code stops working, and it is the difference between a provider that changed something and a provider that broke something. ## Go further - [Errors](/docs/platform/conventions/errors) - [Requests and responses](/docs/platform/conventions/requests-and-responses) --- ## ARNs and identifiers Source: https://shelfcs.com/docs/platform/conventions/arns-and-identifiers The ARN grammar, the partition, account ids, resource ids, and where the rules differ ## Objective Identifiers are the least reversible decision in the platform. They appear in policies, in metering records, in audit records, in error messages, in customer scripts and in infrastructure state files, from the first request onwards. **Read this page before implementing anything that mints an identifier.** ## Requirements - None ## Instructions ### Resource names ``` arn:::::/ ``` ``` arn:aws-shelf:ec2:hel1:123456789012:instance/i-0a1b2c3d4e5f60718 arn:aws-shelf:iam::123456789012:user/deploy arn:aws-shelf:billing:hel1:123456789012:invoice/2026-000123 ``` The region segment is empty for global services. The account segment is present except for resources owned by the platform rather than by a customer. ### The partition **PROPOSED:** `aws-shelf`. The partition cannot be an arbitrary word. The Terraform AWS provider validates partitions against `^aws(-[a-z]+)*$`, and it constructs ARNs itself — in `aws_iam_policy_document`, in the `aws_arn` data source, and in every resource that synthesises one. A partition outside that pattern is rejected by the provider, and the ARNs the provider generates would not match ours. **Additionally, and permanently, ARNs arriving with the `aws` partition are accepted on input.** Terraform derives the partition from provider metadata and will generate `arn:aws:` for a region it does not recognise, so a policy written through it would otherwise never match anything. This is the least reversible decision in the document set: it is embedded in every stored policy and every Terraform state file the moment a customer writes one. ### Resource-type separator The separator between resource type and resource id is a slash for every resource type currently defined. Where a future service mirrors an AWS service that uses the colon form — `resource-type:id` — that service documents the colon form for its own resources. The grammar permits both; the platform does not silently change the form of an existing resource type. ### Account ids **PROPOSED:** twelve decimal digits, never beginning with zero. See [Accounts and tenancy](/docs/platform/conventions/accounts-and-tenancy). ### Infrastructure resource ids **PROPOSED:** a type prefix, a hyphen, and 17 lowercase hexadecimal characters. ``` i-0a1b2c3d4e5f60718 ``` Seventeen hexadecimal characters because that is the modern AWS long-id format, and because customer tooling validates against it. The Terraform AWS provider, among others, validates instance ids against `^i-([0-9a-f]{8}|[0-9a-f]{17})$`. Seventeen hex is safe; anything else is not. | Prefix | Resource | | --- | --- | | `i-` | Instance | | `ami-` | Image | | `vol-` | Volume | | `snap-` | Snapshot | | `key-` | Key pair | | `eni-` | Network interface | | `sg-` | Security group | | `vpc-` | Network | | `subnet-` | Subnet | | `quo-` | Quote | | `ord-` | Order | Prefixes for resources that do not exist in v1 are reserved here so that they cannot later be assigned to something else. These ids are opaque. Customers must not parse them, and nothing in one encodes the account, the region, the creation time, or anything about the hardware. **PROPOSED:** ids are generated from a cryptographically secure source, so that one id reveals nothing about the existence, count or ordering of others. ### Billing document numbers are different Invoice numbers are **not** opaque and **not** random: ``` 2026-000123 ``` They are sequential, gapless within a sequence, and human-readable, because tax authorities require gapless sequential numbering of invoices. An opaque random identifier would not satisfy that. This is a deliberate exception to the opacity rule rather than an inconsistency, and it applies only to documents that are legal records: invoices and credit notes. Quotes and orders, which are not legal records, use opaque ids. **OPEN (Michael):** the exact invoice numbering scheme, and whether the sequence is per-account or global. ### Customer-chosen names Where a resource carries a name the customer chooses — a key pair name, a policy name — the constraints are documented on that resource and are never narrowed after release. **PROPOSED:** where an AWS counterpart exists, its constraints are adopted exactly. Tooling validates names client-side before calling, so a narrower rule would reject names the client believes are valid, and a wider one would accept names the client refuses to send. ### Case Service names, region names and resource ids are lowercase, and comparison is case-sensitive. Two exceptions, both real: - **Access key ids** are uppercase and case-sensitive. A verifier that lowercases them before lookup breaks every signature. - **HTTP header names** are matched case-insensitively by HTTP, and are lowercased only when building the canonical request for a signature. ## Go further - [Accounts and tenancy](/docs/platform/conventions/accounts-and-tenancy) - [Authentication](/docs/platform/conventions/authentication) --- ## Authentication Source: https://shelfcs.com/docs/platform/conventions/authentication SigV4 signing, credential scope, clock skew, and the rules a verifier must follow ## Objective Every request to every Shelf Cloud API is authenticated. There are no anonymous endpoints. **Read this page to implement a client by hand, to implement the server side, or to diagnose a signature failure.** ## Requirements - An access key id and secret access key ## Instructions ### Signature Version 4 Requests are signed with AWS Signature Version 4 — the same algorithm, not an approximation of it, so that existing clients and SDKs sign correctly without modification. A signed request carries: | Header | Contents | | --- | --- | | `Authorization` | Algorithm, credential, signed headers, signature | | `X-Amz-Date` | The signing timestamp, ISO 8601 basic | | `X-Amz-Content-Sha256` | Required by some services; see below | ### The Authorization header ``` Authorization: AWS4-HMAC-SHA256 Credential=AKIAIOSFODNN7EXAMPLE/20260903/hel1/ec2/aws4_request, SignedHeaders=content-type;host;x-amz-date, Signature=5d672d79c15b13162d9279b0855cfba... ``` The `Credential` parameter is the access key id, a slash, then the credential scope: ``` ////aws4_request ``` **The access key id is not part of the credential scope.** It is sent in the clear as the first element of `Credential`; the scope is the four elements after it. A verifier that splits `Credential` into the wrong number of fields, or that feeds the key id into the signing key derivation, produces signatures no SDK can match. ### The credential scope ``` 20260903/hel1/ec2/aws4_request ``` The scope binds a signature to a date, a region and a service. A signature produced for `ec2` in `hel1` cannot be replayed against `iam`, against another region, or on another day. This is why the endpoint shape — one host per service per region — cannot change once customers are signing against it. ### Signing names Each service documents its signing name. It equals the first label of the endpoint host unless the service's page states otherwise. That qualification is deliberate: AWS's own signing names diverge from host prefixes for several services, and stating the rule as a universal would make it wrong the first time we add a service whose counterpart diverges — by which point it is inside every signature. ### Global services A service that is not regional still needs a region in the scope. **PROPOSED:** global services are signed with `hel1`. **PROPOSED:** the verifier additionally accepts a signature scoped to the caller's own configured region for global services. SDK behaviour when an endpoint is overridden varies between languages and versions, and a customer should not receive `SignatureDoesNotMatch` for a difference we can absorb. ### The payload hash The hash of the payload is the final line of the canonical request **for every SigV4 request, without exception.** A request with no body signs the hash of the empty string. What is service-specific is the `X-Amz-Content-Sha256` *header*, not the hashing. Its absence never means the body is unsigned. A verifier that skips payload hashing for services that do not send that header leaves request bodies unauthenticated: an intermediary could alter the body and the signature would still validate. That is a remote tampering hole, and it is the reason this paragraph is stated as emphatically as it is. ### Which headers must be signed `SignedHeaders` must contain, and the verifier must require: - `host` - `x-amz-date` - every `x-amz-*` header actually present on the request - `content-type`, when a body is sent Headers not listed in `SignedHeaders` are **ignored** by the server rather than acted on. A verifier that acts on an unsigned header lets an intermediary change the meaning of a signed request. The list above is closed. A vaguer rule — "refuse if any meaningful header is unsigned" — is not implementable, because the server cannot know which headers the client considered meaningful, and two implementers would draw the line differently. ### Credentials | Element | Notes | | --- | --- | | Access key id | Sent in the clear. **PROPOSED:** 20 characters, uppercase alphanumeric, with a fixed prefix identifying the credential type | | Secret access key | Never transmitted. **PROPOSED:** 40 characters of base64 | Access key ids are case-sensitive and uppercase. They are the one identifier in the platform that is not lowercase, and a verifier that lowercases them before lookup breaks every signature. The secret is returned once, at creation, and cannot be retrieved afterwards. ### Temporary credentials **OPEN (Michael):** whether temporary credentials exist in v1. Until they do, a request carrying `X-Amz-Security-Token` is **rejected** with `InvalidClientTokenId` rather than ignored. Accepting the signature while discarding the session token would accept a credential the caller believed was time-limited and scoped — worse than refusing it. This is not a corner case. The default AWS credential chain supplies a session token in many environments, including assumed roles and SSO, so a customer may send one without intending to. **PROPOSED:** when temporary credentials do exist, `X-Amz-Security-Token` is part of the canonical request. AWS services differ on whether it is signed or appended after signing; we pick one and document it, because a client guessing the other way produces a mismatch every time. ### Presigned requests **OPEN (Michael):** whether query-string SigV4 — `X-Amz-Algorithm`, `X-Amz-Credential`, `X-Amz-Date`, `X-Amz-Expires`, `X-Amz-SignedHeaders`, `X-Amz-Signature` — is supported. If it is not supported, a request carrying those parameters must be **rejected explicitly** rather than ignored, so that a customer generating a presigned URL learns immediately rather than discovering an unauthenticated path. ### Clock skew **PROPOSED:** 15 minutes, matching what existing clients expect. A request outside the window is rejected with `RequestExpired` and HTTP 400. Our clock is in the `Date` header of every response, including errors. The AWS SDKs use this to correct themselves automatically, but only when the status and the error code are the ones they recognise — which is why the statuses on the [errors page](/docs/platform/conventions/errors) are not open to preference. ### Replay The skew window is the only bound on replay: a captured signed request can be resent within it by anyone who obtains it — from a log, a proxy, or a trace. For reads this is disclosure. For creates that accept a client token the damage is bounded, since a replay returns the original result. For creates without one, a replay creates a second billable resource. **OPEN (Michael):** whether to maintain a nonce cache to reject exact replays within the skew window, and at what cost. ### Rules a verifier must follow Stated here because each has produced a real vulnerability in a shipped implementation of this exact protocol: 1. **Verify possession of the secret.** Validating that a signature is well-formed, or that the scope parses, is not authentication. The signature must be recomputed from the derived signing key and compared. 2. **Compare in constant time.** A byte-by-byte comparison that returns early leaks the signature through timing. 3. **Reject any algorithm other than `AWS4-HMAC-SHA256`.** An implementation that falls through on an unrecognised algorithm is an authentication bypass. 4. **Resolve `Host` deliberately.** Behind a load balancer, decide once whether the signature is checked against `Host` or a forwarded header, and never trust a client-supplied forwarded header — otherwise the host component of the signature is attacker-controlled. 5. **Do not distinguish failures.** Wrong secret, wrong region and unknown key return the same error. ### What authentication does not decide Authentication establishes who is calling. Authorisation is evaluated separately, against policy. A correctly signed request may still be refused. ## Go further - [Errors](/docs/platform/conventions/errors) - [ARNs and identifiers](/docs/platform/conventions/arns-and-identifiers) - Signing requests --- ## Endpoints and regions Source: https://shelfcs.com/docs/platform/conventions/endpoints-and-regions How endpoints are constructed, what regions and zones mean, and which services are global. ## Objective The endpoint shape is embedded in every signature. It cannot be changed after the first customer signs a request. **Read this page before configuring a client, and before implementing a service.** > [!primary] > > This page states the intended convention. For **the hostnames that are > actually published today** — eleven of them, all flat on `shelfcs.com` with no > region label — see > [Service endpoints](/docs/platform/endpoints/service-endpoints). Where the two > disagree, that page describes reality and this one describes the intent. ## Instructions ### Endpoints One host per service, per region: ``` .. ``` ``` ec2.hel1.shelfcloud.com billing.hel1.shelfcloud.com ``` Global services use a single host with no region label: ``` iam.shelfcloud.com account.shelfcloud.com ``` **OPEN (Michael):** the public API domain. There is no endpoint discovery service. Endpoints are constructed from the service name and the region by the rule above, and the rule is documented so that clients can construct them. ### Why the shape is fixed The service name and the region are part of the credential scope, and therefore part of every signature. A client that signed against one host cannot be redirected to a differently-shaped one without every signature failing. This is also why a single global `endpoint_url` override in a client configuration is wrong: it would send every service's calls to one host, and those calls are signed for the service they were meant for. ### Regions A region is a geographic location with its own endpoints, its own resources and its own capacity. Resources do not cross regions, and a request signed for one region is not valid in another. ``` hel1 ``` **PROPOSED:** region names are a location code and a digit, where the code names the facility's location and the digit distinguishes facilities in one location. **OPEN (Michael):** whether to adopt AWS-shaped region names — `eu-north-1` rather than `hel1`. SDKs accept arbitrary strings, so `hel1` is not a hard break, but region-parsing helpers, partition inference in third-party tooling, and anything that regexes AWS region names will misbehave on a name outside that shape. This is worth a deliberate decision rather than an accident. ### Zones A zone is a subdivision of a region, named as the region followed by a letter: ``` hel1-a ``` Every region has at least one zone. A region with one zone says so plainly; it is never implied that there are more. Resources that live in a zone report their zone from the first release, even while a region has only one, so that a caller storing the zone today keeps working when a second exists. ### Which services are regional | Service | Scope | Signing region | | --- | --- | --- | | EC2 | Regional | The region in the host | | Billing | Regional | The region in the host | | IAM | Global | **PROPOSED:** `hel1` | | Account | Global | **PROPOSED:** `hel1` | A global service still requires a region in the credential scope. If the server expects one string and the client signs with another, every call returns `SignatureDoesNotMatch` with no diagnosable cause — so the verifier additionally accepts a signature scoped to the caller's own configured region for global services. **OPEN (Michael):** confirmation of the global/regional split. Moving a service from regional to global later invalidates every stored credential-scope expectation in customer tooling, so it cannot be revisited quietly. **OPEN (Michael):** whether an account with resources in several regions receives one invoice or several, which decides whether Billing stays regional. ### the cloud platform version floor The layer requires the compute service microversion **2.83** (Ussuri) on every deployment it fronts. The floor is platform-wide and stated once, here; action pages cite it rather than each declaring their own. 2.83 is set by `DescribeInstances`: its filter push-down uses `vm_state`, `key_name` and `availability_zone` as non-admin query parameters, which the compute service silently strips below 2.83 rather than rejecting. Every other v1 action needs 2.67 or lower. At start-up the layer reads the compute service's version document (`GET /`) and refuses to serve with `ServiceUnavailable` if the maximum reported is below the floor. An operator on Stein or Train learns this from the start-up log, not from a customer's failing call. **OPEN (Michael):** whether to hold the floor at 2.83, or drop the `DescribeInstances` push-down to keep 2.67 and reach older deployments. ### Transport HTTPS only. HTTP requests are refused, never redirected — a redirect would invite a client to send a signed request over an unencrypted connection first. **PROPOSED:** TLS 1.2 minimum, TLS 1.3 preferred. ## Go further - [Authentication](/docs/platform/conventions/authentication) - [ARNs and identifiers](/docs/platform/conventions/arns-and-identifiers) --- ## Errors Source: https://shelfcs.com/docs/platform/conventions/errors Error envelopes, HTTP status codes, shared codes, and what an error must never reveal ## Objective An error response is part of the contract. Callers branch on error codes, so a code whose meaning changes breaks working software as surely as a removed endpoint does. **Read this page before implementing any service, and before writing any client error handling.** ## Requirements - None ## Instructions ### Two envelopes, not one Shelf Cloud services use one of two error envelopes, decided by the protocol the service speaks. They are not interchangeable, and a client written for one cannot parse the other. **EC2 query envelope** — used by services mirroring Amazon EC2: ```xml InvalidInstanceID.NotFound The instance ID 'i-0a1b2c3d4e5f60718' does not exist b1e2c3d4-5678-90ab-cdef-1234567890ab ``` No `` element. The request id element is `RequestID`. **Standard query envelope** — used by services mirroring IAM and STS: ```xml Sender NoSuchEntity The user with name deploy cannot be found b1e2c3d4-5678-90ab-cdef-1234567890ab ``` Has ``. The request id element is `RequestId`, with different casing from the EC2 form. **JSON envelope** — used by services with no AWS counterpart, such as Billing and Account: ```json { "__type": "QuoteExpired", "message": "Quote quo-0a1b2c3d4e5f60718 expired at 2026-09-03T14:00:00Z" } ``` with `x-amzn-ErrorType` also carrying the code as a response header. | Service | Envelope | | --- | --- | | EC2 | EC2 query | | IAM | Standard query | | Account | JSON | | Billing | JSON | The casing differences between `RequestID` and `RequestId` are inherited and are preserved exactly. A client that parses one and not the other must behave here as it does against AWS. ### Request ids in headers Every response, successful or not, carries the request id in the `x-amzn-RequestId` header. This is the header every AWS SDK reads to populate its response metadata; a custom header would be invisible to all of them, and the customer told to "quote the request id" would have none to quote. Successful EC2-protocol responses additionally carry a `` element in the body, lowercase initial, as Amazon EC2 does. ### HTTP status codes | Status | Meaning | | --- | --- | | `400` | Malformed, invalid, expired, or a resource that does not exist | | `403` | Credentials unknown, or policy denies the action | | `404` | JSON-protocol services only, for a missing entity | | `409` | The request conflicts with the state of the resource | | `500` | An error on our side | | `503` | Temporarily unavailable, or the request rate was exceeded | **There is no `401` anywhere in this platform.** AWS uses none in this protocol, and two things depend on that: - The AWS SDKs correct their own clocks automatically when they see a 403 or 400 carrying a recognised skew-related code together with our `Date` header. A 401 defeats that recovery entirely. - RFC 9110 requires a 401 response to carry a `WWW-Authenticate` header. We send none, so a 401 would be a protocol violation, and intermediaries that special-case 401 would behave unpredictably. ### Shared codes Returned by every service. | Code | Status | Meaning | Retry | | --- | --- | --- | --- | | `IncompleteSignature` | 400 | The `Authorization` header is malformed | No | | `InvalidAction` | 400 | Unknown action, or an unimplemented version | No | | `InvalidParameterValue` | 400 | A parameter has an unacceptable value | No | | `InvalidParameterCombination` | 400 | Parameters cannot be used together | No | | `MissingParameter` | 400 | A required parameter is absent | No | | `InvalidPaginationToken` | 400 | The token is invalid or expired | No | | `IdempotentParameterMismatch` | 400 | Token reused with different parameters | No | | `RequestExpired` | 400 | The timestamp is outside the permitted skew | After fixing the clock | | `AuthFailure` | 403 | The credentials could not be validated | No | | `SignatureDoesNotMatch` | 403 | The signature does not match the request | No | | `InvalidClientTokenId` | 403 | The access key id is unknown or inactive | No | | `MissingAuthenticationToken` | 403 | The request was not signed | No | | `UnauthorizedOperation` | 403 | Policy denies this action on this resource | No | | `AccessDenied` | 403 | Policy denies this action | No | | `InternalError` | 500 | An unexpected failure on our side | Yes | | `ServiceUnavailable` | 503 | Temporarily unavailable | Yes | Note the statuses carefully. `InvalidClientTokenId` is **403**, not 401. `RequestExpired` is **400**, not 401 and not 403. `IncompleteSignature` is **400** while `SignatureDoesNotMatch` is **403**. These are AWS's own values, and matching them is what makes existing client recovery logic work. ### Throttling differs by service | Service kind | Code | Status | | --- | --- | --- | | EC2-mirroring | `RequestLimitExceeded` | 503 | | IAM-mirroring | `Throttling` | 400 | | JSON services | `Throttling` | 429 | This is not an inconsistency we chose. Amazon EC2 returns `RequestLimitExceeded`, and the Terraform AWS provider matches that exact string for its backoff; a 429 there would produce no automatic retry at all for a large installed base of clients. Services with no AWS counterpart are free to use the modern 429 and do. Each service's error page states which applies to it. ### Absence versus denial A caller must never learn that a resource exists by the shape of the error. The rule is a decision procedure, not a preference: 1. **Another account's resource, or no such resource** → the service's not-found code. Identical responses in both cases. 2. **This account, resource exists, policy denies** → `UnauthorizedOperation` or `AccessDenied`. So `UnauthorizedOperation` confirms existence — and it may only be returned within the caller's own account, where existence is not a secret from them. ### What an error never contains No stack traces. No internal hostnames. No database or query text. No detail about another account. No hint that distinguishes a wrong password from an unknown user, or a wrong signature from a wrong key. `SignatureDoesNotMatch` is deliberately identical whether the secret was wrong, the region was wrong, or the clock was wrong. An error that distinguished them would help an attacker as much as a customer, and the troubleshooting guides carry the differential diagnosis instead. ### Messages The `Code` is stable and is what callers branch on. The `Message` is written for a human, may change wording between releases, and must never be parsed. ### Deprecating a code A code is never given a new meaning. A behaviour needing a different meaning gets a new code, and the old code keeps returning what it always did until it is removed with notice. ## Go further - [Making requests](/docs/platform/conventions/requests-and-responses) - [Authentication](/docs/platform/conventions/authentication) --- ## Idempotency Source: https://shelfcs.com/docs/platform/conventions/idempotency Client tokens, which operations accept them, and how a retry is made safe ## Objective Networks fail after the server has acted and before the client has heard. Anything that creates or destroys must therefore be safe to retry, or a lost response becomes a duplicate resource and a duplicate charge. **Read this page before writing any automated caller.** ## Requirements - None ## Instructions ### Client tokens An operation that accepts a `ClientToken` takes a value the caller chooses, unique to that logical request. If a request arrives with a token already seen: - **and the parameters are identical**, the original result is returned and nothing new is created; - **and the parameters differ**, the request is refused with `IdempotentParameterMismatch`. The mismatch case exists so that a bug in a caller — reusing a token while changing what it asks for — surfaces as an error rather than as a silent duplicate. **PROPOSED:** tokens are up to 64 ASCII characters and are remembered for 24 hours. After that a repeated token is treated as new. **PROPOSED:** a UUID is the recommended token. The documentation says so rather than leaving each caller to invent a scheme. ### Which operations accept a token **Only those whose mirrored AWS model declares one.** This is a hard constraint rather than a style choice. The AWS SDKs validate parameters against their own bundled service model *before sending anything*. A token added to an operation that AWS does not model with one is rejected by the client with a parameter-validation error, and the request never leaves the customer's machine. Mandating a token everywhere would mandate something customers cannot send. `RunInstances` accepts one. `CreateSecurityGroup`, `CreateKeyPair`, `CreateTags` and `CreateVolume` do not, because Amazon EC2's 2016-11-15 model does not declare one for them. Services with no AWS counterpart — Billing, Account — are free to require a token, and `PlaceOrder` does require one, because it is the operation that spends money. Each operation's reference page states which case it is in. ### Where an operation has no token Retry safety comes from the operation's own semantics, and each reference page says which applies: - **Absolute writes.** `CreateTags` sets tags to the values given; repeating it changes nothing. - **Deletes.** See below. - **Neither.** `CreateKeyPair` generates a new key pair each time it is called. A retry after a lost response produces a second key pair, which is untidy but not billable. The reference page says so plainly. ### Deletes A delete returns the mirrored AWS error when the resource does not exist. It does not report success. This is worth stating precisely because the opposite convention is tempting and wrong. Infrastructure tooling — the Terraform AWS provider among others — polls for a not-found error to confirm that a destroy has completed, and treats not-found during a refresh as "deleted out of band". A server that always reported success for a delete would leave that tooling unable to distinguish "gone" from "not yet gone", producing hung destroys and permanent state drift. Retry safety for deletes therefore belongs to the client: a caller retrying a delete treats the not-found error as success. That is what the SDKs and provider already do. ### Interaction with billing An idempotent replay never produces a second billable resource and never opens a second usage stream. Usage records are written against the resource over time, not against the request that created it, so a replayed request that creates nothing meters nothing. ### Reads Always safe, always repeatable. A read is never rate-limited into failure by being repeated, only throttled. ## Go further - [Errors](/docs/platform/conventions/errors) - [Making requests](/docs/platform/conventions/requests-and-responses) - [Metering events](/docs/platform/conventions/metering-events) --- ## Metering events Source: https://shelfcs.com/docs/platform/conventions/metering-events The usage record, its fields, and its interval semantics. ## Objective Metering is a platform convention rather than an implementation detail of Billing, because usage history cannot be reconstructed after the fact. Anything not recorded correctly on the first day cannot be billed for later and cannot be defended in a dispute. **Read this page before any service emits its first usage record.** ## Requirements - None ## Instructions ### The record Every billable thing emits usage records containing at least: | Field | Type | Meaning | | --- | --- | --- | | `RecordId` | string | Unique; the deduplication key | | `AccountId` | string | Who is billed | | `ResourceArn` | string | What was used | | `Region` | string | Where, stated rather than parsed out of the ARN | | `MeterName` | string | Which meter, from a fixed vocabulary | | `Quantity` | integer | How much; may be negative on a correction | | `Unit` | string | The unit of the quantity | | `StartTime` | timestamp | Inclusive | | `EndTime` | timestamp | Exclusive | | `PriceVersion` | string | Which catalogue version applied | | `CorrectsRecordId` | string | Present only on a correcting record | `Region` is carried explicitly rather than derived from the ARN. Global services have an empty region segment in their ARNs, and parsing an identifier to recover billing information contradicts the rule that identifiers are opaque. `PriceVersion` exists so that an invoice can be recomputed and defended after a price change. Without it, an old invoice cannot be reproduced from its inputs. ### Interval semantics The interval is half-open: **`[StartTime, EndTime)`**. The start instant is included; the end instant is not. Consecutive records for one resource therefore abut exactly. No second is counted twice, and no second is lost, at every boundary between every pair of records. Left ambiguous, two engineers would choose differently and the error would be systematic across millions of records rather than occasional. ### When a record is written A record is written **for the interval in which the usage occurred**, as it occurs — not at creation of the resource, and not when the invoice is prepared. A resource that exists over time produces a series of records covering that time. A single record written at creation could not express a duration, and nothing that is billed by elapsed time could be invoiced at all. An idempotent API replay creates no new resource and therefore opens no second series of records. ### Resolution **PROPOSED:** usage is recorded at one-second resolution. Recording is deliberately finer than charging: the charging rule can then be stated exactly rather than approximated, and a finer record can always be aggregated while a coarser one can never be refined. **OPEN (Michael):** the charged granularity, and the rounding rule for a partial unit — what is rounded, at which step, and in which direction. It must appear both here and in the Billing user guide, because a customer reconciling an invoice needs it. ### The clock Usage timestamps come from the platform, not from inside a customer's instance, so that a clock set wrongly inside an instance cannot affect what is charged. **OPEN (Michael):** which component is authoritative, and what happens to records from a host whose clock is found to be wrong. The 15-minute skew tolerance that applies to API requests does not apply here: in metering, a clock error is money. ### Meters The meter vocabulary is fixed and versioned like an API. A meter name is never reused for a different measurement, and a meter whose definition must change gets a new name. **OPEN (Michael):** the v1 meter vocabulary. ### Corrections Records are immutable. A record found to be wrong is never edited. A correction is a new record that names the record it corrects in `CorrectsRecordId` and carries the difference in `Quantity`, which may be negative. The total usage of a resource is therefore the sum of its records, corrections included, and the history of what was believed at each point remains readable. Without the linking field, a correction is indistinguishable from additional usage and an invoice cannot be explained line by line. ### Retention **OPEN (Michael):** how long raw usage records are retained. This is a legal question as much as a technical one and belongs in the terms as well as here. Note that request logs are proposed at 90 days. If usage records are retained for longer — and tax law will require that they are — a dispute about an old period can be resolved from usage records but not from request logs. That is acceptable, but it should be a decision rather than a surprise. ### Egress **Egress is recorded and is not charged for.** The distinction is the whole point. Not charging is a product commitment, and it can be revisited by a future release without breaking anything. Not *recording* would be irreversible: no history to bill from later, and no visibility of abuse until it arrives as an invoice from an upstream transit provider. So egress produces usage records like everything else. No egress line appears on any invoice. **OPEN (Michael):** confirmation that egress metering is built in v1 even though nothing is charged for it. ## Go further - [Time, numbers and money](/docs/platform/conventions/time-numbers-and-money) - [Idempotency](/docs/platform/conventions/idempotency) - Usage and metering --- ## Pagination Source: https://shelfcs.com/docs/platform/conventions/pagination Which operations paginate, token behaviour, and the empty-page trap ## Objective **Read this page before writing any loop over a list operation.** The most common integration bug against this platform is stopping on an empty page. ## Requirements - None ## Instructions ### Which operations paginate **Only those whose mirrored AWS model declares `MaxResults` and `NextToken`.** This is a constraint rather than a preference. The AWS SDKs validate parameters against their own service model before sending, and their paginators exist only for operations the model marks as paginated. Two failures follow from getting this wrong: - Adding `MaxResults` to an operation AWS does not model with one means the customer's SDK rejects the call before it is sent. - Returning a `NextToken` from an operation AWS models as unpaginated means the SDK **discards it silently**. The caller receives a truncated list, with no error, and believes it is complete. Half the regions, half the zones, and no indication that anything is missing. `DescribeInstances`, `DescribeVolumes`, `DescribeSnapshots` and `DescribeImages` paginate. `DescribeRegions`, `DescribeAvailabilityZones` and `DescribeAccountAttributes` do not, and must never return a token. Services with no AWS counterpart paginate wherever a list can grow. ### Parameters | Parameter | Type | Notes | | --- | --- | --- | | `MaxResults` | integer | Items per page | | `NextToken` | string | From a previous response | **PROPOSED:** `MaxResults` accepts 5 to 1000, defaulting to 1000, for EC2-mirroring services — matching Amazon EC2's own bounds. `MaxResults` may not be combined with an explicit list of resource ids. That combination returns `InvalidParameterCombination`, as it does at Amazon EC2. ### The token `NextToken` is opaque. Do not parse it, construct it, or store it beyond the sequence it belongs to. Its format may change without notice; nothing about it is documented, because anything documented becomes something a caller depends on. **A response that omits `NextToken` is the last page. That is the only signal that a listing has ended.** A page may contain fewer items than `MaxResults`, or none at all, and still carry a token. This happens whenever filters remove every item from a page, and it is normal. ```python token = None while True: kwargs = {"MaxResults": 1000} if token: kwargs["NextToken"] = token page = client.describe_instances(**kwargs) handle(page["Reservations"]) token = page.get("NextToken") if not token: # correct break # if not page["Reservations"]: break ← wrong; silently loses results ``` **PROPOSED:** a token is valid for 24 hours, after which `InvalidPaginationToken`. ### Filters Filters are applied **per page**, after the page is assembled — not before pagination. This matches Amazon EC2, and it is why an empty page with a token occurs at all. The alternative, filtering before paginating, would require scanning the entire result set to fill each page. There is no guarantee that a filtered page is non-empty. ### Ordering **PROPOSED:** results are ordered by creation time, oldest first, stable across pages. ### Consistency A listing is not a snapshot. Items created during a pagination sequence may or may not appear. Items deleted during one may still appear, and a later read of one may return not-found. A caller needing a consistent view must reconcile after the listing completes rather than assuming the listing provided one. ## Go further - [Requests and responses](/docs/platform/conventions/requests-and-responses) - [Errors](/docs/platform/conventions/errors) --- ## Requests and responses Source: https://shelfcs.com/docs/platform/conventions/requests-and-responses Request structure, headers, request ids, retries, consistency and throttling ## Objective **Read this page to construct a request by hand, or to understand what your SDK is doing for you.** ## Requirements - An access key id and secret access key ## Instructions ### Protocol HTTPS, HTTP/1.1 or HTTP/2. Each service uses the wire protocol its AWS counterpart uses, because matching it is what makes existing clients work: | Service | Protocol | Request | Response | | --- | --- | --- | --- | | EC2 | Query | Form-encoded `POST`, or `GET` | XML | | IAM | Query | Form-encoded `POST` | XML | | Account | JSON | JSON `POST`, action in `X-Shelf-Target` | JSON | | Billing | JSON | JSON `POST`, action in `X-Shelf-Target` | JSON | There is no content negotiation. A service speaks one protocol. ### Required request headers | Header | Purpose | | --- | --- | | `Authorization` | The SigV4 signature | | `X-Amz-Date` | Request timestamp, ISO 8601 basic | | `Host` | Signed; must match the endpoint | | `Content-Type` | As the service requires | ### Response headers | Header | Purpose | | --- | --- | | `x-amzn-RequestId` | The request id, on every response | | `Date` | Our clock, for skew detection | | `x-amzn-ErrorType` | The error code, on JSON-service errors | `x-amzn-RequestId` is the header the AWS SDKs read to populate response metadata. It is emitted on success and on failure, including on authentication failures. ### Request ids Every request has an id. It is the only thing support needs to find a call in our records, and it is returned in the header above and in the response body for query-protocol services. **PROPOSED:** request ids are retained for 90 days. That retention is shorter than the retention of usage records, deliberately, and the difference has a consequence worth stating: a billing dispute raised about a period older than 90 days can be resolved from usage records but not from request logs. ### Sizes **OPEN (Michael):** maximum request body size, header size, and URL length. ### Compression **PROPOSED:** `gzip` accepted on requests and offered on responses when the client advertises it. ### Redirects Never. An endpoint answers or returns an error. A redirect would invite a client to re-send a signed request to a host it did not intend to call, and the host is part of the signature. ### Retries Retry on 500, 502, 503 and 504, and on the throttling code for that service, with exponential backoff and jitter. Do not retry a 4xx — **except** the eventual-consistency case below. ### Eventual consistency A resource that has just been created may not be visible to an immediately following call. A describe issued moments after a create may return a not-found error for a resource that does exist. This is inherited from Amazon EC2, whose own documentation instructs callers to allow for it, and existing tooling already does: the Terraform AWS provider retries `InvalidInstanceID.NotFound`, `InvalidGroup.NotFound` and `InvalidVpcID.NotFound` for exactly this reason. **So: retry a not-found error for a resource created within the last few seconds.** This is the one documented exception to the rule that a 4xx is not retryable, and a client that follows the general rule without this exception will fail intermittently under load. **PROPOSED:** a created resource is visible to all readers within 10 seconds. **OPEN (Michael):** whether any operation offers read-after-write consistency, and if so which. Every service that does not must carry a propagation note on its create operations. ### Throttling Each service documents its own throttling code and status; they differ, and the [errors page](/docs/platform/conventions/errors) explains why. **OPEN (Michael):** per-account request rates, per service. A throttled response carries a distinct code rather than a generic error, so that a client can distinguish "slow down" from "this will never work". ## Go further - [Errors](/docs/platform/conventions/errors) - [Authentication](/docs/platform/conventions/authentication) - [Idempotency](/docs/platform/conventions/idempotency) --- ## Time, numbers and money Source: https://shelfcs.com/docs/platform/conventions/time-numbers-and-money Timestamp formats, units, and the exact representation of money ## Objective **Read this page before writing any code that parses a Shelf Cloud response.** ## Requirements - None ## Instructions ### Time All timestamps are UTC, ISO 8601 extended, with a `Z` suffix. Services mirroring an AWS API emit **millisecond** precision, because their AWS counterparts do and clients parse accordingly: ``` 2026-09-03T14:22:07.000Z ``` Services with no AWS counterpart emit second precision unless a field documents otherwise: ``` 2026-09-03T14:22:07Z ``` No local times, no offsets, no epoch seconds in any customer-facing field. A field carrying a date without a time is documented as a date. Durations are integers of seconds, and the field name says so. Our clock is in the `Date` header of every response so that a client can detect its own drift before signatures begin failing. ### Numbers Integers are integers. Nothing that counts is returned as a string. Sizes are in bytes unless the field name carries another unit, and where another unit is used the name carries it: `SizeGiB`, `MemoryMiB`, `ThroughputMbps`. Binary units are binary — a GiB is 1024³. Decimal units are decimal — a Mbps is 10⁶ bits per second. The two are never mixed inside one field, and the field name always says which is meant. ### Money Monetary amounts are an integer of minor units with an explicit currency: ```json { "Amount": 1980, "Currency": "EUR" } ``` That is €19.80. Never a float, never a decimal string. Floating point cannot represent decimal money exactly, and a rounding error in an invoice is not a rounding error, it is a dispute. **OPEN (Michael):** the billing currency, and whether more than one is offered. A currency is per-account and permanent, so it must be settled before any account is created. ### Rates A rate needs more precision than a price, because a per-second rate is a small fraction of a minor unit: ```json { "Amount": 550, "Scale": 6, "Currency": "EUR", "Unit": "instance-second" } ``` `Scale` is the number of decimal places beyond the minor unit that `Amount` carries. The example is 0.000550 cents per instance-second. **OPEN (Michael):** the rounding rule — at which step rounding occurs, and in which direction. Rounding per record, per line and per invoice give different totals, and the difference is visible to any customer who checks. ### Net and gross Amounts are stated **net of tax** unless a field name says otherwise. Tax is a separate line, and the gross total is stated separately. | Field | Meaning | | --- | --- | | `AmountNet` | Before tax | | `TaxAmount` | Tax on that base | | `AmountGross` | Net plus tax | An amount with no statement of whether it includes tax is not a reconcilable figure, and for a provider selling across EU borders it is a legal exposure rather than an inconvenience. **OPEN (Michael):** the tax model — rate determination, place of supply, the reverse charge for VAT-registered business customers in other member states, and where the VAT identification number is captured. ### Percentages and ratios Decimals with a documented scale, or basis points where a field says so. Never a string with a percent sign. ### Null and absent A field that is absent is absent. A field present and null means the value is known to be nothing. Empty strings are never used to mean absent. ## Go further - [Metering events](/docs/platform/conventions/metering-events) - Billing API data types --- # Platform — ENDPOINTS ## Service endpoints Source: https://shelfcs.com/docs/platform/endpoints/service-endpoints The eleven hostnames this cloud publishes, what speaks each one, and how they differ from the endpoint shape the conventions page proposes. ## Objective **These are the hosts a customer can actually reach today.** The [Endpoints and regions](/docs/platform/conventions/endpoints-and-regions) convention proposes a `..` shape with the domain marked `OPEN (Michael)`; what is deployed is flatter and the domain is decided. Where the two disagree, this page describes reality and the convention page describes the intent. One list, three consumers: the deployment registers these hostnames in the identity service catalog, routes them to the services behind them, and declares their DNS. ## The hosts All on `shelfcs.com`, all HTTPS on 443. | Hostname | Service | the cloud platform project | Local port | | --- | --- | --- | --- | | `identity.shelfcs.com` | Identity | the identity service | 5000 | | `ec2.shelfcs.com` | EC2-compatible API | ec2-api | 8788 | | `compute.shelfcs.com` | Compute | the compute service | 8774 | | `image.shelfcs.com` | Images | the image service | 9292 | | `network.shelfcs.com` | Networking | the networking service | 9696 | | `volume.shelfcs.com` | Block storage | the block-storage service | 8776 | | `placement.shelfcs.com` | the placement service | the placement service | 8780 | | `rating.shelfcs.com` | Rating | the rating service | 8889 | | `console.shelfcs.com` | Web console | the web console | 9999 | | `console-api.shelfcs.com` | Web console API | the web console API | 9998 | | `vnc.shelfcs.com` | Instance console proxy | noVNC | 6080 | There is **no endpoint discovery service**. Construct a host from this table, or read the identity service catalog with `--os-interface public`. ## How the traffic reaches us ``` client → edge (TLS terminates here) → private transport → the service ``` Every API binds to a private network only; no service port is exposed to the internet directly. Traffic reaches them through a single managed ingress, which means: - There is no origin address to attack. - TLS terminates at the edge. - **If the ingress is down, every endpoint is down at once.** There is no second path and no failover. > [!warning] > > **The ingress is not yet doing everything it should.** There is no rate > limiting configured on it, and none on the authentication endpoint. Until > those exist, treat the absence as a gap rather than as permission. ## Regions and the credential scope There is one region, `hel1` (Hetzner Helsinki), and one zone, `hel1-a`. The region name is **not** in the hostname, but it **is** in the SigV4 credential scope. A client must be told `hel1` explicitly: ``` aws --endpoint-url https://ec2.shelfcs.com --region hel1 ec2 describe-instances ``` ```hcl provider "aws" { region = "hel1" skip_credentials_validation = true skip_requesting_account_id = true skip_metadata_api_check = true endpoints { ec2 = "https://ec2.shelfcs.com" } } ``` Get the region wrong and every call returns `SignatureDoesNotMatch` with nothing in the message that tells you why. That is SigV4's behaviour, not ours. `hel1` is not an AWS-shaped region name, and anything that regexes AWS region names will not like it — whether to move to `eu-north-1` is `OPEN (Michael)` on the conventions page. ## Where the EC2 endpoint differs from the rest `ec2.shelfcs.com` is the only host that speaks **SigV4 with an access key pair**. Signatures are verified by the identity service's the identity service’s signature-verification call, so an EC2 access key is the identity service credential underneath. Every other host on the table speaks **the identity service tokens** — `X-Auth-Token`, obtained from `identity.shelfcs.com`. See [Account developer guide](/docs/account/guide/developer-guide). Two credentials, one identity. Revoking the identity service user revokes both. ## What is not published - **No object storage.** There is no Swift, no S3-compatible endpoint. Anything written for S3 has nothing to talk to here. - **No DNS, no load balancer (Octavia), no managed Kubernetes, no managed database, no queue, no secrets service, no telemetry API.** - **No `billing.` or `account.` host.** The conventions page names both; what exists instead is the storefront Worker — see [Account API](/docs/account/cloud/storefront-api-reference). - **No status page.** An unpublished service is not a service running behind a firewall waiting to be switched on. It does not exist. ## Go further - [Endpoints and regions](/docs/platform/conventions/endpoints-and-regions) — the proposed convention - [Authentication](/docs/platform/conventions/authentication) - [Account developer guide](/docs/account/guide/developer-guide) - [Web console and VNC](/docs/platform/console/web-console-and-vnc) --- # Platform — SECURITY ## Credentials and trust Source: https://shelfcs.com/docs/platform/security/credentials-and-trust Which of your secrets this platform can hand back to you, what that means for blast radius, and where the real trust boundaries are. ## Objective **Read this before deciding how much to trust a session on this platform.** Most clouds are built so that a secret, once issued, cannot be retrieved — only replaced. This one is not, in two places, and both are deliberate design choices with consequences you should know about rather than discover. ## Your cloud password is stored, and it is recoverable The cloud password you set at onboarding is kept in the accounts database, in a column that holds it as text — not a hash. It is kept for a working reason: the console signs in with a username and a password and nothing else, so storing it is what lets the account page open the console for you without asking you to type it again. ``` POST /v1/console/session → { "username": "cust-…", "password": "…", "region": "…", "console_url": "…" } ``` That endpoint hands your own password back to your own session. > [!warning] > > **A password kept in recoverable form cannot be treated like a hashed one.** > > - Anyone who obtains a session token for your site account can read your cloud > password, not merely act on your behalf while the token lasts. > - **Do not reuse this password anywhere else.** Not your email, not your other > clouds, not anything. Treat it as a value the platform holds, because it is. > - There is no password-reset endpoint for it yet, so you cannot rotate it > yourself if you think it has leaked. That is the gap that matters most here. ## Access key secrets are also recoverable `GET /v1/access-keys` returns the **secret**, not just the key id. Amazon shows a secret access key exactly once because it does not keep it in a form it could show again; the identity service here does keep it, and this API returns it. The consequence is the same shape as above: a session token is not a limited credential. It is a route to every long-lived secret the account holds. Full detail: [Access keys](/docs/account/cloud/access-keys). ## So what is the real blast radius | If this leaks | An attacker gets | | --- | --- | | A site session token (15 min) | Your cloud password, every access key secret, your billing portal link — **everything**, and permanently, because the secrets outlive the token | | One access key pair | Full control of your project — no key is narrower than the user | | Your cloud password | Both consoles and every cloud API | **There is no credential on this platform that is less powerful than any other.** No read-only key, no scoped token, no expiry, no MFA, no condition keys. That is documented in [Identity, as deployed](/docs/identity/keystone/identity-as-deployed); this page is about what follows from it. Practically: the site session is the crown jewel. Protect it accordingly, and do not leave one open on a shared machine. ## What is not stored - **Your site login password.** That belongs to the hosted auth service and this platform never sees it. - **Card details.** They belong to the payment provider; the platform holds a customer reference, not a card. - **The contents of your machines or volumes.** Nothing reads them. ## Where the trust boundaries actually are | Boundary | What it separates | | --- | --- | | Your account | You from every other customer. **The only boundary you get** | | Your project's network and router | Your machines from other customers' machines | | Security groups | Your own tiers from each other — the only tool inside your account | | The private network | Nothing from you: every machine you run can reach every other | Two things that must not touch each other need **two accounts**. ## What is not protected yet Stated so nobody assumes otherwise: - **No rate limiting.** The public edge has none configured, and neither does the authentication endpoint. Do not read the absence as permission — it will arrive. - **No MFA**, on either identity system. - **No audit log you can read.** Services log server-side; nothing surfaces it to you, so you cannot see when your own credentials were used. - **No IMDSv2.** The SSRF defence it exists for is unavailable — see [Instance metadata and user data](/docs/compute/ec2/ec2-instance-metadata-and-user-data). - **No encryption at rest for volumes.** `Encrypted` is refused rather than faked, so at least a module asking for it fails loudly. - **No provider backups.** Nothing here protects you from losing data, only from someone else reading it. ## What to do about it 1. **A unique cloud password**, stored in a password manager, reused nowhere. 2. **One access key per consumer** — CI, laptop, tooling — so withdrawing one does not stop the others. They are not narrower, but they are individually revocable. 3. **Separate accounts** for environments that must not reach each other. 4. **Assume a leaked session is a full compromise**, and rotate every key if one happens. The cloud password you cannot yet rotate; tell us instead. 5. **Do not put the only copy of anything here.** ## Go further - [Identity, as deployed](/docs/identity/keystone/identity-as-deployed) - [Access keys](/docs/account/cloud/access-keys) - [Authentication](/docs/platform/conventions/authentication) - [Account developer guide](/docs/account/guide/developer-guide) --- ## What we verify before selling it Source: https://shelfcs.com/docs/platform/security/verified-claims Three checks that must pass before this platform will serve — the hardware is as big as the catalog claims, every alias tells the truth about size, and no allocation oversells the box. ## Objective Most of what a cloud tells you about itself is a claim. Some of what this one tells you is **checked by a machine, and the deployment refuses to proceed if the check fails**. This page says which claims those are, so you know which ones are load-bearing. ## 1. The hardware is at least as big as the catalog sells Before the platform serves, it measures the machine it is running on and compares it with what the catalog offers: | Resource | Rule | | --- | --- | | **CPU** | Real threads × the declared overcommit ceiling must cover the advertised vCPU | | **Memory** | Physical memory must cover the advertised memory. **Never overcommitted** | | **Storage** | Free pool space plus what has already been sold must cover the advertised capacity | If any of the three falls short the deploy **fails**, with a message naming the resource, the real figure and the catalog figure. The consequence for you: the capacity figures on [Instance types](/docs/compute/ec2/ec2-api-instance-types) are not aspirational. The platform will not run on hardware that cannot honour them. Memory is the one with no give in it. CPU may be oversubscribed up to the ceiling the catalog itself declares; memory never is, so a shape promising 8 GB has 8 GB of real memory behind it. ## 2. Every AWS alias tells the truth about size Our shapes carry AWS instance-type names as aliases, so existing tooling works unchanged — `m5.large` resolves to a shape of ours. **That claim is verified, not asserted.** On every deploy, each alias is checked against a public feed of real AWS instance specifications: - An alias **unknown to AWS** fails the deploy. - An alias whose **vCPU or memory differs** from the shape it points at fails the deploy, with a message naming both sides — *"`x` is 4c/16384MiB on AWS but 2c/8192MiB in `cd-…`"*. The spec feed is refreshed daily and is the same data the pricing rule uses, so the size an alias promises and the price derived from it cannot drift apart. > [!primary] > > **An alias is never a lie about size.** That is why shapes with no honest AWS > equivalent carry no alias at all rather than a near-enough one — the 1-vCPU > shapes have none, because no current-generation, non-burstable AWS type is > that small. What this buys you: `instance_type = "m5.large"` in a module written for AWS gets exactly the vCPU and memory `m5.large` means. It does not get "roughly that". ## 3. No set of allocations can oversell the box Every tenancy ceiling is declared, and the **first** thing the process that creates them does is add them all up and assert the total fits inside what the catalog says the hardware sells. Adding a tenancy, or raising an existing one, past that line **fails before anything is touched** — no partial application, no first-come-first-served race discovered later under load. The fix is never to raise the ceiling. It is to shrink another allocation, or to free capacity. See [Quotas and limits](/docs/account/cloud/quotas-and-limits). ## What these checks do not cover Stated so the guarantees are not read as broader than they are: - **Not availability.** Nothing here promises the machine stays up. One box, no redundancy, no failover. - **Not performance under contention.** The volume IOPS figure is published and divided honestly, but it is **not yet enforced** per volume — see [Volume types](/docs/storage/block/volume-types). - **Not durability.** No backups, no mirroring across disks, no cross-host replication. - **Not price stability over time.** Prices are recomputed from a published rule; the rule is fixed, the numbers move with the market. - **Not that everything on offer is priced.** A shape with no rating mapping is launchable and unbilled — a defect, not an offer. ## Why publish this Because the interesting question about any cloud's specification is not what it says, but which parts of it something would notice if they became untrue. Here, three things would: the hardware being smaller than the catalog, an alias lying about size, and allocations summing past capacity. Everything else on this platform is a claim you are trusting us on, and the pages describing it try to say so plainly. ## Go further - [Instance types](/docs/compute/ec2/ec2-api-instance-types) — the shapes and their aliases - [How prices are set](/docs/billing/pricing/how-prices-are-set) — the rule the same feed drives - [Quotas and limits](/docs/account/cloud/quotas-and-limits) — the ceilings that are summed - [Volume types](/docs/storage/block/volume-types) — a figure that is published but not yet enforced --- # STORAGE — BLOCK ## Volume types Source: https://shelfcs.com/docs/storage/block/volume-types The one block-storage type we sell, the performance it delivers, how that figure was measured, and where it is not yet enforced. ## Objective **Read this page before choosing a volume size or building anything latency-sensitive on this platform.** It says what the disk underneath a volume actually is. the only source of truth for the offer; this page restates it and nothing else. ## The type | Field | Value | | --- | --- | | Our name | `hdd` | | EC2 name (`VolumeType` on the wire) | `standard` | | Minimum size | 1 GiB | | Maximum size | 1000 GiB | | IOPS | 48 | | Throughput | 11 MB/s | | Default | Yes — the only type | There is one type because the box has one kind of disk: **7200 rpm SATA**. No SSD, no NVMe. `standard` is Amazon's own name for magnetic storage, so the alias is honest. There is deliberately no `gp3` or `io2` alias: those names promise SSD latency this hardware cannot deliver, and a customer whose module says `type = "gp3"` should get an error, not a slow disk that looks like a fast one. ## Where 48 IOPS comes from It is measured, then divided. | Step | Value | | --- | --- | | Measured with `fio` on 2026-08-23, whole pool | 580 random-read IOPS, 664 random-write IOPS, 132 MB/s sequential | | Most instances the box can hold | 24 vCPU ÷ 2 vCPU per instance = 12 | | Per-volume figure published | 580 ÷ 12 = 48 IOPS, 132 ÷ 12 = 11 MB/s | The division is the point. **A figure quoted to one customer is a figure every customer can have at once.** The alternative — publishing the whole-pool number and hoping only one customer is busy — is how a shared HDD box becomes an outage. ## What that means in practice 48 IOPS is a spinning disk's share, and it is small. Concretely: - **A database on a volume will be slow**, and the slowness will be seek time, not bandwidth. Postgres, MySQL and anything else doing small random writes are the worst fit for this storage. - **Sequential work is fine.** Backups, logs, media, build artefacts, object data, anything that streams — 11 MB/s sustained per volume, more in bursts while the pool is quiet. - **Put the working set in RAM.** The memory shapes (`cd-memory-*`) exist because on this hardware buying RAM is cheaper than buying IOPS. If your workload needs SSD latency, this is the wrong platform today, and we would rather say so on this page than discover it together on your first production incident. > [!warning] > > **The IOPS figure is not enforced yet.** The platform's first rule is that > every capability figure is enforced, not advertised — if the catalog says 48 > IOPS, the hypervisor limits the disk to 48 IOPS. For volumes that limit does > not exist today: `bootstrap/roles/openstack/tasks/resources.yml` creates the compute service > flavors from the catalog and nothing else, so there is no the block-storage service QoS > specification associated with the `hdd` type. > > Until that task exists, one busy volume can take the whole pool's throughput > and every other volume on the box slows down. This is a known defect. It is > tracked as our internal readiness list gap 5 and as open decision 2 on the > [Volume actions](/docs/compute/ec2/ec2-api-volumes) page. ## Capacity, and what shares the pool Every volume, and every instance root disk, is carved from one LVM volume group on one machine. | Figure | Value | Source | | --- | --- | --- | | Pool name | `default` | | Sellable capacity | 1400 GiB | | Physical | 1.7 TB, ZFS | | Reserved for instance root disks | 1200 GiB, thin LV `nova-instances` | | Volume GiB already sold | Tracked internally — free space plus this must still cover the catalog | Root volumes are thin-provisioned. **The sum of what has been sold is capped**, so the pool cannot fill silently behind customers who are each within their own quota. A create that would take the pool past the cap is refused. ## Root disk against attached volume They are not the same thing and they are not priced the same way. | | Root disk | Attached volume | | --- | --- | --- | | Where it comes from | The instance type's `root_gib` (see [Instance types](/docs/compute/ec2/ec2-api-instance-types)) | `CreateVolume` | | Size | Fixed by the shape, 20–160 GiB | 1–1000 GiB, growable | | Survives termination | No — `deleteOnTermination` is `true` | Yes | | Cost | Included in the instance price | **Not billed today** — see below | | Backend | the compute service's thin LV `nova-instances` | the block-storage service, same pool | ## Durability, honestly - One box. One pool. **The pool is not mirrored across disks today** (our internal readiness list gap 1: one HDD holds the OS and the only pool; the second disk is not yet attached as a mirror). - There is no cross-host replication and no cross-region copy. - Snapshots are stored in the same pool as the volumes they protect. A snapshot protects you against `rm -rf`, not against the disk. - Backups with a tested restore are our internal readiness list gap 2 and do not exist. Treat a volume as durable against a process crash or a guest mistake, not against a hardware failure. Take snapshots, and copy anything you cannot lose off the platform. ## Pricing Attached storage is **not rated today**. The catalog's pricing rule prices the AWS-equivalent instance type and nothing else, which is why the storage-dense shape `cd-storage-4-16-1000` was removed from the catalog rather than sold at a compute price. The practical consequence: a volume you create is free, and will not stay free. See [How prices are set](/docs/billing/pricing/how-prices-are-set). ## Go further - [Volume actions](/docs/compute/ec2/ec2-api-volumes) — the EC2 API for volumes - [Storage developer guide](/docs/storage/guide/developer-guide) — the block-storage service directly - [Snapshot actions](/docs/compute/ec2/ec2-api-snapshots) - [Instance types](/docs/compute/ec2/ec2-api-instance-types) --- # STORAGE — GUIDE ## Storage developer guide Source: https://shelfcs.com/docs/storage/guide/developer-guide Working with the volume API — the calls in order, the states you must wait for, the retry trap that bills you twice, and code for the whole attach-format-mount cycle. ## What this is Developer-focused information about driving block storage from code: which call to make, what state a volume must be in before it will accept it, what to retry and what never to retry, and the guest-side half that no API can do for you. Concepts are in the [user guide](/docs/storage/guide/user-guide); every parameter is in the [API reference](/docs/compute/ec2/ec2-api-volumes). ## Endpoint Volume actions are EC2 actions and go to the compute endpoint, not to a storage one: ``` https://ec2.shelfcs.com region hel1 ``` The same volumes are reachable through the block-storage service's own API on its own host — see [Storage developer guide](/docs/storage/guide/developer-guide). Both views are of one object; only the identifier differs. **Manage a volume through the API you created it with.** ## The state machine ``` CreateVolume → creating → available ─ AttachVolume → in-use ↖ DetachVolume ↙ available ─ DeleteVolume → deleting → gone ``` Every call has a state precondition, and violating it is an error rather than a wait: | Call | Volume must be | Instance must be | | --- | --- | --- | | `AttachVolume` | `available` | `running` or `stopped`, same zone | | `DetachVolume` | `in-use` | — | | `ModifyVolume` | `available` or `in-use` | — | | `DeleteVolume` | `available` | — | So `create → attach` immediately fails: the volume is `creating` for a moment first. Poll `DescribeVolumes` until `status` is `available`, or use a waiter. ```python ec2.get_waiter("volume_available").wait(VolumeIds=[vol]) ec2.attach_volume(VolumeId=vol, InstanceId=inst, Device="/dev/sdf") ec2.get_waiter("volume_in_use").wait(VolumeIds=[vol]) ``` ## The retry trap > [!warning] > > **`CreateVolume` is not idempotent and cannot be made so.** The Amazon EC2 > model declares no `ClientToken` for it, so no SDK will send one. A retry after > a lost response creates a **second volume**, which occupies the pool and — once > storage is billed — bills for it. Do not wrap `CreateVolume` in a blind retry. Instead: 1. Tag at creation: `--tag-specifications 'ResourceType=volume,Tags=[{Key=Name,Value=data-01}]'` 2. On retry, `DescribeVolumes --filters Name=tag:Name,Values=data-01` first. 3. Create only if nothing came back. Terraform does this for you — `aws_ebs_volume` records the id in state before it retries anything. Everything else is safe: `AttachVolume`, `DetachVolume` and `DeleteVolume` are naturally idempotent, and a delete of an already-deleted volume returns `InvalidVolume.NotFound`, which every client treats as success. ## Parameters that are refused, not ignored | Parameter | Result | | --- | --- | | `VolumeType` other than `standard` | `InvalidParameterValue` | | `Iops`, `Throughput` | `InvalidParameterValue` | | `Encrypted`, `KmsKeyId` | `InvalidParameterValue` | | `MultiAttachEnabled` | `InvalidParameterValue` | | `Size` outside 1–1000 | `InvalidParameterValue` | The AWS provider defaults `type` to `gp2`, so **omitting `type` fails at apply**. Write `type = "standard"` explicitly. Refusing rather than ignoring is deliberate: a module that sets `encrypted = true` by policy fails loudly instead of shipping unencrypted storage that reports success. ## Growing a volume `ModifyVolume` changes the volume. It does **not** change the filesystem. ```bash sc ec2 modify-volume --volume-id vol-… --size 40 ``` `DescribeVolumesModifications` is **not implemented** — poll `DescribeVolumes` for the new `size` instead. A client that waits on the modifications call gets `InvalidAction`. Then, inside the guest: ```bash # an attached volume needs a rescan before the kernel sees the new size echo 1 > /sys/class/block/vdb/device/rescan growpart /dev/vdb 1 resize2fs /dev/vdb1 # or xfs_growfs /data ``` Until you do that, you are paying for space the guest cannot see. And a `Size` at or below the current one is an error, not a no-op — a silent no-op would make a shrink attempt look like it worked. ## The guest-side half No API call formats or mounts anything. The full cycle: ```bash lsblk # find it — the device name may not be what you asked for mkfs.ext4 /dev/vdb mkdir -p /data blkid /dev/vdb # take the UUID echo 'UUID= /data ext4 defaults,nofail 0 2' >> /etc/fstab mount -a ``` Automate it in user data, and make it re-runnable — check for a filesystem before making one, or a re-run wipes the disk: ```yaml #cloud-config runcmd: - [ bash, -c, 'blkid /dev/vdb || mkfs.ext4 /dev/vdb' ] - [ bash, -c, 'grep -q /data /etc/fstab || echo "UUID=$(blkid -s UUID -o value /dev/vdb) /data ext4 defaults,nofail 0 2" >> /etc/fstab' ] - [ mount, -a ] ``` Remember that `runcmd` swallows failures — see [Instance metadata and user data](/docs/compute/ec2/ec2-instance-metadata-and-user-data). ## Detaching safely ```bash umount /data # inside the guest, FIRST sc ec2 detach-volume --volume-id vol-… ``` `--force` exists for an unresponsive machine and risks filesystem corruption. It does not stop the instance first. Detaching does **not** stop billing. The volume exists until you delete it. ## Errors | Code | Cause | Do | | --- | --- | --- | | `VolumeInUse` | Attached already, or you tried to delete an attached volume | Detach first | | `IncorrectState` | Wrong state for the call | Wait, then retry | | `InvalidVolume.ZoneMismatch` | Volume and instance in different zones | There is one zone — check what you passed | | `VolumeLimitExceeded` | Your quota | Ask for more | | `InsufficientVolumeCapacity` | **The pool is full**, not your quota | Nothing you can change | | `InvalidVolume.NotFound` | No such volume, or another account's | Treat as not-found | `InsufficientVolumeCapacity` is the one to handle differently: retrying, or asking for quota, will not help. ## Terraform ```hcl resource "aws_ebs_volume" "data" { availability_zone = "hel1-a" size = 20 type = "standard" # required; the gp2 default is rejected tags = { Name = "data" } } resource "aws_volume_attachment" "data" { device_name = "/dev/sdf" volume_id = aws_ebs_volume.data.id instance_id = aws_instance.app.id skip_destroy = true # when the volume outlives the instance force_detach = false # leave it } ``` `aws_ebs_volume` with `size` reduced is a **replace**, not a shrink — Terraform will destroy the volume and its data. Guard anything you care about with `lifecycle { prevent_destroy = true }`. ## Go further - [User guide](/docs/storage/guide/user-guide) - [API reference](/docs/compute/ec2/ec2-api-volumes) - [Volume types](/docs/storage/block/volume-types) - [Storage developer guide](/docs/storage/guide/developer-guide) - [Idempotency](/docs/platform/conventions/idempotency) --- ## Storage user guide Source: https://shelfcs.com/docs/storage/guide/user-guide Getting started with block storage — what a volume is, root disk against attached volume, sizing for the hardware you actually have, and what you must do yourself. ## What this is A **volume** is a block device you attach to one machine. It exists independently of any instance, it survives termination, and it grows but never shrinks. There is no object storage, no file share, no managed database. Block storage is the only storage service, and this guide is how to use it well on the hardware it actually runs on. Then: [developer guide](/docs/storage/guide/developer-guide) for the API mechanics, [API reference](/docs/compute/ec2/ec2-api-volumes) for every operation. ## Two kinds of disk, and only one survives | | Root disk | Attached volume | | --- | --- | --- | | Where it comes from | The instance shape, 20–160 GB | You create it, 1–1000 GB | | Survives terminate | **No** — deleted with the instance | **Yes** | | Resizable | No | Yes, upward only | | How many per machine | One | Several | | Cost | Included in the instance price | Not billed today | **Anything you want to keep goes on an attached volume.** The root disk is scratch space with an operating system on it. ## What the disk actually is One type, because the box has one kind of disk: **7200 rpm SATA**. No SSD, no NVMe. | | | | --- | --- | | Our name | `hdd` | | Name on the wire | `standard` | | Size | 1–1000 GB | | IOPS | 48 | | Throughput | 11 MB/s | Those figures are **measured and then divided**: `fio` on the whole pool gave 580 random-read IOPS and 132 MB/s, divided by the twelve machines the box can hold at once. So the number quoted to you is one every customer can have simultaneously, not a best case that evaporates when a neighbour wakes up. There is deliberately no `gp3` or `io2` alias. Those names promise SSD latency this hardware cannot deliver, and a module that asks for `gp3` gets an error rather than a slow disk that looks like a fast one. ## Sizing for spinning disks 48 IOPS is a small number and it is seek-bound, not bandwidth-bound. | Workload | Verdict | | --- | --- | | Logs, backups, media, build artefacts, anything sequential | **Good.** 11 MB/s sustained per volume | | A web app's static files and code | **Fine** | | Postgres, MySQL, anything doing small random writes | **Bad.** This is the worst fit for this storage | | Search indexes, heavy analytics | **Bad** | The way to win here is to **buy memory instead of IOPS**: the `cd-memory-*` shapes exist because on this hardware RAM is the cheaper way to make a working set fast. Put the hot data in page cache and the disk stops mattering. If your workload genuinely needs SSD latency, this is the wrong platform, and we would rather you read that here than discover it during an incident. ## What you must do yourself - **Format and mount.** Attaching does not partition, format or mount anything. - **Grow the filesystem.** Growing the volume does not grow what is on it. - **Unmount before detaching.** Detaching a mounted, written filesystem is pulling the disk out. - **Take your own snapshots — and delete them.** Nothing is scheduled, and nothing expires them. They occupy the same pool as your volumes, so a forgotten nightly snapshot eventually fills it. - **Copy anything irreplaceable off the platform.** There are no provider backups, and snapshots live in the same pool as the volume they protect. ## Durability, plainly One box. One pool. The pool is **not mirrored across disks** today. There is no cross-host replication and no cross-region copy. A volume is durable against a process crash, a bad deploy or an `rm -rf` you took a snapshot before. It is **not** durable against the disk failing. Plan accordingly, and do not let this platform hold the only copy of anything. ## Identify disks by UUID, never by device name The device name you ask for on attach is a request, not a guarantee — you may ask for `/dev/sdf` and the guest kernel may call it `/dev/vdb`. A reboot can renumber devices, and an `/etc/fstab` written against a path will then mount the wrong disk. ```bash blkid /dev/vdb echo 'UUID= /data ext4 defaults,nofail 0 2' >> /etc/fstab ``` `nofail` is not optional. Without it, a machine whose volume is missing at boot drops into emergency mode wanting a root password that a cloud image does not have — and recovering means detaching the root disk and mounting it elsewhere. ## Capacity you share Every volume and every instance root disk comes out of one 1400 GB pool, of which 1200 GB is reserved for root disks. Volumes are thin-provisioned but the sum sold is capped, so the pool cannot fill silently behind customers who are each within their quota. A create that would take the pool past the cap fails with `InsufficientVolumeCapacity` — capacity, not policy. A quota increase will not fix it. Snapshots come out of the same pool. See [Snapshot actions](/docs/compute/ec2/ec2-api-snapshots). ## Two things that are not yet true > [!warning] > > **The IOPS figure is not enforced.** The platform's rule is that every > capability figure is enforced rather than advertised; for volumes that limit > does not exist yet, so one busy volume can take the pool's throughput. This is > a known defect. > > **Storage is not quota'd and not billed.** Onboarding sets CPU and memory > quotas only, and no volume is rated. Neither is a gift — both will change, and > a cost model built on free storage will break. ## Go further - [Developer guide](/docs/storage/guide/developer-guide) — the API mechanics - [API reference](/docs/compute/ec2/ec2-api-volumes) — every operation - [Volume types](/docs/storage/block/volume-types) — the measurement in full - [Snapshot actions](/docs/compute/ec2/ec2-api-snapshots) - [Compute user guide](/docs/compute/guide/user-guide) ---