Object storage¶
nilo_s3 reads and writes objects in a bucket — S3, MinIO, R2, Backblaze,
anything that speaks the same dialect. A bucket is a type, and a key is
not (ADR 059): the
bucket's name is compiled in, so a handler asks for it by type the way it
asks for a database, and the key is data the way a path param is.
It is a Service: it borrows the event loop and holds a named destination —
an endpoint, a region, credentials. SigV4 and S3's semantics are all it is;
the HTTP underneath is nilo_fetch, which makes this the only
module that imports a Fitting
(ADR 058,
ADR 063).
const s3 = @import("nilo_s3");
and in build.zig, beside nilo_http:
.{ .name = "nilo_s3", .module = nilo.module("nilo_s3") },
The whole of it¶
const s3 = @import("nilo_s3");
const Avatars = s3.Bucket("avatars", .{ .max_bytes = 2 << 20 });
fn avatar(avatars: *Avatars, c: *nilo.Ctx, key: nilo.Str) !void {
const object = try avatars.get(c, key.view());
return c.send(200, object.content_type.view(), object.bytes.view());
}
and at startup:
var store = try s3.open(gpa, .{
.endpoint = "https://s3.ap-southeast-1.amazonaws.com",
.region = "ap-southeast-1",
.credentials = .{ .static = .{
.access_key_id = settings.aws_key,
.secret_access_key = settings.aws_secret,
} },
});
defer store.deinit();
var avatars = try Avatars.open(&store);
defer avatars.deinit();
try app.provide(&avatars);
One Store, many Buckets. The Store holds what changes between a laptop and production — the endpoint, the credentials, the signing key derived from them, and one connection pool. A Bucket holds a name and the options that belong to the bucket itself. Two buckets over one Store are two types, two services a handler can name, and one pool.
open dials nothing. A Service that needs the loop is finished when the loop
exists (ADR 037),
so the first credential fetch and the first connection happen inside
listen(), before anything is accepted. A bucket asked for before that
answers error.Failed, with the reason in the log, rather than dialling
nothing.
c is a Scope — the *Ctx, or a nilo.Run for work
outside a request.
Reading¶
| Call | |
|---|---|
bucket.get(c, key) |
Object — bytes, content_type, etag, len, all Str in the Scope, from one allocation |
bucket.getRange(c, key, .{ .from = 0, .to = 1023 }) |
the same, for a slice. to is inclusive, the way HTTP counts |
bucket.getIf(c, key, etag) |
Conditional — .unmodified or .{ .object = … }. A 304 is a success, so it is a union rather than an error |
bucket.head(c, key) |
Meta — len, content_type, etag, and no bytes |
bucket.stream(c, key, &reading) |
the object left on the socket: below |
bucket.list(c, .{ .prefix, .max_keys, .cursor }) |
Page — objects and next: one page of keys, and where the next page starts. Below |
max_bytes is checked against content-length before a byte is read,
so an object over it costs one round trip rather than a download and comes
back as error.TooLarge. The bucket's ceiling is the largest thing a
bounded get will ever put in a request arena, which is why it is on the
type: it is a fact about what the program is prepared to hold.
getIf is what an endpoint that forwards an ETag wants — the browser's
If-None-Match becomes the bucket's, and .unmodified becomes a 304 with no
body moved at either end.
Listing what is there¶
var cursor: ?[]const u8 = null;
while (true) {
const page = try files.list(c, .{ .prefix = "exports/2026/", .max_keys = 200, .cursor = cursor });
for (page.objects) |o| {
// o.key, o.size, o.etag, o.last_modified — every text a Str in the Scope
}
cursor = (page.next orelse break).view();
}
A list is the one call here that asks about the bucket rather than about a
key, and it is bounded three ways on purpose
(ADR 058).
A page is at most max_keys objects, and at most 1,000, which is S3's own
ceiling — asking for more is error.Rejected rather than a page quietly
capped, so a loop sized by what it asked for is never given less. The body is
read into the Scope up to a ceiling worked out from max_keys and the
bucket's key_max, so a server answering more than the question is
error.TooLarge rather than an arena it fills. And nothing follows the
cursor for you: next is handed back as text, the loop above is yours, and
there is no listAll. A helper that walked every page would be the call with
unbounded output the module does not have, and the one that turns a bucket
into a database.
etag comes back quoted, the way head and get hand it back, so it can go
straight into getIf. last_modified is the server's own text —
2026-09-18T10:11:12.000Z — and sql.Timestamp reads it if you need
arithmetic. What it costs is two allocations for the body and the page, and
one per key and one per ETag for the decoding: a key arrives percent-encoded
(the request asks for it that way, so a key with an & in it is not an XML
problem) and an ETag arrives with its quotes as entities.
Writing¶
| Call | |
|---|---|
bucket.put(c, key, .{ .bytes = …, .content_type = … }) |
also reads .cache_control and .content_disposition if the value has them |
bucket.putStream(c, key, .{ .reader = …, .len = …, .content_type = … }) |
from a *std.Io.Reader of a known length |
bucket.delete(c, key) |
put takes anything with .bytes and .content_type, checked while
compiling — which is the shape a nilo.Upload already
has, so a file out of a form goes straight through:
const NewAvatar = struct { image: nilo.Upload };
fn setAvatar(avatars: *Avatars, c: *nilo.Ctx, account: u32, incoming: nilo.Form(NewAvatar)) !nilo.Status(201, void) {
var key: [32]u8 = undefined;
try avatars.put(c, try std.fmt.bufPrint(&key, "{d}.png", .{account}), incoming.value.image);
return .{};
}
The key is yours, and image.filename is not it: that is whatever the
browser sent, ../../etc/passwd included. Encoding it into a URL is
nilo_s3's job and is done once; deciding what an object is called is the
application's.
putStream frames by length, never chunked, because S3 answers 411 to
a body whose length it was not told. The length is a field rather than an
optional so that I do not know it is a compile error here rather than S3's
status code. An upload whose size is unknown before it starts is a multipart
upload, which is not built.
An object too big to hold¶
get puts the object in the request arena, which is right for an avatar and
wrong for a video. stream leaves it on the socket:
const s3 = @import("nilo_s3");
const Videos = s3.Bucket("videos", .{});
fn watch(videos: *Videos, c: *nilo.Ctx, key: nilo.Str) !void {
var reading: Videos.Reading = .idle;
defer reading.close();
try videos.stream(c, key.view(), &reading);
var body = try c.streamWith(200, reading.content_type, .{ .length = reading.len });
_ = try reading.pipe(&body.writer);
try body.finish();
}
Reading carries len, content_type and etag once stream returns,
then pipe(w) moves the bytes into any *std.Io.Writer — a response, a
file — and close() gives the connection back. content_type and etag
are borrowed and valid only until pipe: they point into the connection's
read buffer, which the first byte of body reads over, so set the response
headers from them first and copy them if you want them afterwards. And a
Reading must not be copied once begun, for the reason a fetch.Exchange
must not: it holds one.
.length = reading.len is what gives the browser a Content-Length to draw
a progress bar against and a Range to resume with
(Streaming).
There is no buffer to declare. Until 0.5 stream took one, and the
guide said a bigger one was fewer trips into the connection; the bytes go
from the connection's own read buffer to your writer and never crossed it,
which bench/result/s3.md had already seen as a 64 KB to 8 KB change worth
one byte. How much one read brings in is the client's read_buffer_size
(ADR 186).
Letting the browser talk to the bucket¶
Two calls hand somebody else a way in without the bytes passing through your server. Neither opens a socket: both are the signature and nothing else.
| Call | |
|---|---|
bucket.presign(c, key, seconds) |
Presigned — url and expires_at. A link to fetch one object |
bucket.presignPost(c, key, .{ .seconds = 900 }) |
Posted — url, fields and expires_at. A form for uploading one |
expires_at is the true number. A URL signed with temporary credentials
dies when they do, not when X-Amz-Expires says, so the life reported is the
smallest of what was asked for, the bucket's presign_max, and what the
credentials have left. A caller storing it in a database has one that is
right.
A presigned POST is a form rather than a link: url is the bucket, and
fields go into the form in the order they come back, with the file input
last — S3 ignores whatever follows the file part. The first field is
bucket, because Garage refuses the form without it and AWS ignores it.
<form action="{url}" method="post" enctype="multipart/form-data">
<!-- one hidden input per field, in order -->
<input type="file" name="file"> <!-- last -->
</form>
s3.Post |
Default | |
|---|---|---|
seconds |
— | clamped the same three ways presign's are |
content_type |
null |
an eq condition on $Content-Type. Null lets the browser send what it likes |
max_bytes |
the bucket's | clamped to it, and defaulted to it, so a form with no ceiling is not something this hands out |
prefix |
false |
key is the start of a key rather than the whole of one, so the browser picks the filename |
The reason the POST policy is here rather than in your application is one line of SigV4: it is signed with the key nilo derives once a day (ADR 060), and a second implementation outside nilo is two places that have to agree about a rotation. They disagree at 00:00 UTC, and the symptom is uploads failing with a 403 that says nothing.
The Store's options¶
Given to s3.open:
| Field | Default | |
|---|---|---|
endpoint |
— | https://host[:port] or http://host[:port], no path |
public_endpoint |
null |
the endpoint a browser reaches, when it is not the one this process dials — below |
region |
us-east-1 |
|
credentials |
— | .static or .fetch — below |
max_in_flight |
32 | calls at once, across the process. An HTTPS connection holds 59,151 bytes, so this times that is the store's ceiling |
timeout_ms |
30,000 | one call, end to end |
max_drain |
64 KiB | how much of a refused body is worth reading to keep the connection |
refresh_margin_s |
300 | how long before expiry temporary credentials are replaced |
The scheme decides whether payloads are hashed — UNSIGNED-PAYLOAD over
TLS, a real SHA-256 over plaintext — and there is nothing to configure
(ADR 060). It follows
that the plaintext numbers in bench/result/s3.md
carry a hash over every body that the HTTPS ones would not, and the HTTPS
ones carry a record layer the plaintext ones do not.
An endpoint that is not scheme://host[:port] is error.BadEndpoint at
open, and a region longer than a credential scope can carry is
error.BadRegion.
A store the browser reaches another way¶
A process dials the store on a Docker network or a Tailscale address; a
browser holding a presigned URL cannot. public_endpoint is the name the
browser reaches, and it is the host in every presigned URL and every POST
form — signed as that host, which is why rewriting the URL after the
fact cannot do this job: the host is inside the signature, and a presigned
GET with a rewritten host is a 403 that reads like a signing bug
(ADR 177).
Every call that dials — get, put, head, delete — still dials and
signs endpoint.
var store = try s3.open(gpa, .{
.endpoint = "http://garage:3900",
.public_endpoint = "https://files.example.com",
.credentials = .{ .static = .{ .access_key_id = cfg.s3_key, .secret_access_key = cfg.s3_secret } },
});
The Bucket's options¶
The second argument to s3.Bucket. Every field is a property of the bucket
itself; anything that changes between development and production belongs on
the Store.
| Field | Default | |
|---|---|---|
max_bytes |
8 MiB | the largest object a bounded get will hold |
style |
.virtual |
https://avatars.s3…/key. .path — https://s3…/avatars/key — for MinIO and anything on a bare host |
sse |
null |
.aes256 or .aws_kms, sent as the header S3 already understands |
presign_max |
3600 | the longest life a presigned URL from this bucket may claim |
key_max |
512 | the longest key. Comptime because it sizes a stack buffer at 3× this, and stack is per connection; S3's own ceiling is 1,024 |
session_token_max |
0 | room for a session token beside a signature. Zero is right for static credentials and costs nothing; an STS source sets 2048 and pays for it per connection |
A name that could never work — empty, or not a legal DNS label under
.virtual — stops the build with a sentence, which is the reason the name is
a comptime string rather than a field. So does a credential written into the
options: secret, access_key and their relatives are refused by name,
because a bucket's type is compiled into the binary and ships with it, and
credentials belong on the Store where a Config can read them.
Credentials¶
.credentials = .{ .static = .{ .access_key_id = "…", .secret_access_key = "…" } },
is a key pair for the life of the process. Temporary credentials — IRSA, IMDS, anything STS hands out — are a function:
.credentials = .{ .fetch = fromIrsa }, // fn (gpa: std.mem.Allocator, io: std.Io) anyerror!s3.Credentials
called once at startup and again when the ones in hand are within
refresh_margin_s of expiring. Called lazily, by the request that
notices — there is no background task — and five minutes early, so the
request paying for the refresh is never one that would otherwise have failed.
Whatever the function allocates from gpa is freed by the Store when the
next refresh replaces it. A Credentials carries session_token and
expires_at for exactly this case, and the bucket's session_token_max has
to be raised from zero to make room for the token.
What it answers instead¶
Seven errors, because a handler would do something different about each:
| Error | Answer | |
|---|---|---|
error.NotFound |
no object at that key | 404 — the only one with a default, because its meaning does not change with the request |
error.TooLarge |
the object is over max_bytes, refused before a byte was read |
yours |
error.Throttled |
S3 is shedding load — SlowDown, or a 503 |
a default image and a 200, or a 503; only the handler knows |
error.Unavailable |
S3 answered 5xx. Distinct from Throttled because backing off is the answer to one and not the other |
502 or 503 |
error.TimedOut |
this call's deadline | 504 |
error.Rejected |
S3 refused the request — credentials, signature, permissions | 500, not 403: telling the caller they are not allowed when the truth is that the server's credentials are wrong is a lie in the one place it costs the most debugging |
error.Failed |
everything else, including a response nilo could not make sense of | 502 |
S3's own code and message are logged rather than sent on. A skewed clock —
RequestTimeTooSkewed — is read out of the body and said plainly in the
log, because it is the one S3 error whose fix is on your machine.
What it costs¶
Per request, nothing beyond the call: signing is one SHA-256 over the canonical request, one HMAC, and the hex, because the derived key changes once a day and is kept for the day (ADR 060). The canonical request is never assembled as bytes — it is written straight into the hash — which is why the module has no per-request buffer for it.
Per idle connection, what a handler's stack adds: the encoded key at
3 × key_max, and the depth nilo_fetch drives the fiber to.
bench/result/s3.md measures all of it against
the same seven routes written in Go and Rust, and says which route was the
expensive one and why it was not the one anybody expected.
What it will not do¶
COPY and multipart upload, for one reason rather than two: they are where
S3 stops being bytes at a key and starts being a document format. COPY
can answer 200 with an error in the body, which is a client that has to read
XML to know whether it succeeded; multipart is a protocol — initiate, N parts
with their own ETags, a completion document — rather than a call. Both are
on the roadmap with that reason
attached, waiting on a caller. list was the third of these until
ADR 058,
and what let it in is above: five names scanned for, one page, a cursor, and
no helper that follows it.
Arbitrary x-amz-meta-* is refused on a performance argument that can be
revisited with a measurement: a fixed header set is what keeps
SignedHeaders a walk rather than a per-request sort.
Testing¶
zig build test-s3 runs the module's tests against a canned S3 on a
loopback socket that checks every signature, with no container and no
network — and, with S3_ENDPOINT and friends set, against a real one.
The quickest one to develop against is SeaweedFS in a container, which is what CI runs, with bench/s3_setup.py filling it with the bucket and objects the tests and the benchmark want:
docker run -d --name seaweedfs -p 9100:8333 \
-e AWS_ACCESS_KEY_ID=niloadmin -e AWS_SECRET_ACCESS_KEY=nilosecret123 \
chrislusf/seaweedfs:4.47 server -dir=/data -s3 -s3.port=8333
python3 bench/s3_setup.py
Then .endpoint = "http://127.0.0.1:9100" and .style = .path.
A handler that takes a *Avatars is an ordinary function, and the shape to
test is the one that does not need the bucket — what key a request maps to,
what a NotFound becomes — with the bucket call left to the module's own
suite.
See also¶
- The reference — the surface as a list.
- Calling somebody else's API — the client underneath, and the
ExchangeaReadingis built on. - Forms — the
Uploadthatputtakes as it is. - ADR 059 — why the name is compiled in and the key is not.