Network access: http_fetch
pydeno has no permission model. A runtime grants nothing, and what guest code can reach is exactly
what the host bound into it. So there is no --allow-net=host flag; the equivalent is a host tool
that does the fetching and decides what is allowed. pydeno.http_fetch is that tool, written once
with the server-side request forgery (SSRF) cases handled:
from pydeno import AgentSandbox, http_fetch
fetch = http_fetch(
["api.example.com/v1/", "status.example.com"],
headers={"Authorization": "Bearer ..."}, # the guest can neither set nor read these
timeout=10.0,
max_response_bytes=256 * 1024,
)
with AgentSandbox({"fetch_url": fetch}, max_tool_calls=20) as session:
prompt = session.describe_tools() # lists the allowed URLs for the model
result = session.run("""
const r = await fetch_url("https://api.example.com/v1/items?limit=5")
if (r.status !== 200) return { error: r.status }
return JSON.parse(r.body)
""")
The guest passes one URL and gets back plain data:
{ status: 200, headers: { "content-type": "application/json" }, body: "...", truncated: false,
url: "https://api.example.com/v1/items?limit=5" }
statusis the final status after redirects. 404 and 500 are results, not errors.headersholds only the allow-listed response headers (response_headers=, by defaultcontent-type,content-length,content-language,etag,last-modified,cache-control,expires,date,retry-after), lowercased.Set-Cookieand the rest are dropped.bodyis text (decoded with the response charset, else UTF-8, bad bytes replaced) or, withresponse="bytes", raw bytes (aUint8Arrayin JavaScript).truncatedis set when the body was cut atmax_response_bytes, or the server closed the connection before sending the length it announced.urlis the URL that answered, after redirects.
A refused or failed request throws in the guest, with an error name it can branch on:
e.name |
When |
|---|---|
HttpFetchBlocked |
The policy refused it: URL, scheme, host, resolved address, redirect, size of the URL |
HttpFetchTimeout |
The whole call (DNS, connect, TLS, redirects, body) ran past timeout |
HttpFetchFailed |
Allowed, but it failed: no such host, connection refused, TLS or HTTP error |
All three are pydeno.ToolError subclasses (HttpFetchError is their base). Their messages are
fixed strings written by pydeno ("the URL is not in the allow-list"), never a resolved address,
a response body or a header value, so they are shown to the guest even when the session redacts
host errors. A model can read why it was refused and correct its next call.
Binding it
The tool is an ordinary callable, so it goes wherever a tool goes:
| Where | What to pass |
|---|---|
AgentSandbox({"fetch_url": fetch}) |
fetch, or fetch.aio |
AsyncAgentSandbox({"fetch_url": fetch}) |
either; the sync form already runs off the event loop |
ToolBridge({"fetch_url": fetch.aio}).attach(runtime) |
fetch.aio, so a slow request does not block the runtime thread |
fetch.aio is the async form of the same tool: the blocking socket work runs in a worker thread.
Every call is a tool call, so it counts against the session's max_tool_calls (or the bridge's
max_calls), and in an AgentSandbox it is recorded in the journal like any other: its result is
plain data, and a session restored with AgentSandbox.load() replays it from the journal without
issuing the request again.
Options
| Option | Default | Meaning |
|---|---|---|
allow |
required | Exact destinations, see below |
schemes |
("https",) |
Add "http" only deliberately |
headers |
none | Fixed request headers (an API key). Transport headers (Host, Content-Length, Transfer-Encoding, Connection, ...) are refused |
timeout |
10.0 |
Seconds for the whole call, a deadline rather than a per-read timeout |
max_response_bytes |
1 MiB | Body bytes kept; the rest is never read |
max_url_length |
2048 | The request-size cap (GET only, no request body) |
max_redirects |
5 | Hops followed, each checked from scratch |
redirects |
"same-origin" |
"allow-list" follows to any allowed URL; "never" returns the 3xx |
response |
"text" |
Or "bytes" |
response_headers |
see above | Response header names passed back |
resolver |
system | resolver(host, port) -> [ip, ...]; its answers are vetted like the system's |
ssl_context |
ssl.create_default_context() |
For a private CA, say. Keep host name verification on (see the limits below) |
Allow entries are matched exactly:
"api.example.com": any path on the default port of the scheme."api.example.com:8443": that port only."api.example.com/v1/": paths under/v1/(and/v1itself), on a segment boundary, so/v1evildoes not match."https://api.example.com/v1/": that scheme only.
There is no wildcard and no suffix matching: example.com does not allow www.example.com. Hosts
are compared lowercase after IDNA encoding (bücher.example is xn--bcher-kva.example). An IP
address is allowed only in canonical form (203.0.113.7, [2001:db8::1]), and a private or
loopback one is refused at request time anyway. Cloud metadata names (metadata.google.internal,
metadata, instance-data, ...) are refused even if listed.
Threat model
The guest is untrusted code, often written by a model that read untrusted text. It controls the URL and nothing else. The tool runs in the host process, inside the host's network, which can usually reach things the internet cannot: cloud metadata services, admin ports on loopback, databases on the private network. The goal is that the only thing the guest can make the host fetch is a GET to an allow-listed URL on a public address, with a bounded response.
What every request, and every redirect hop, goes through:
- Strict parsing. The URL must be absolute and ASCII outside the host. Whitespace and
control characters are refused outright, so a CR/LF cannot inject a header or a second request
(percent-encoded
%0D%0Astays encoded on the wire). Refused as well: userinfo (https://user:pass@host/,https://allowed@evil/), IPv6 zone ids, a port that is out of range or not the entry's, and anything in the path that two servers could read differently and use to step out of a path prefix:./..segments (also percent-encoded, and in a relative redirect before it is joined), encoded/or\,;path parameters (/v1/..;/adminis/adminto backends that strip them),%25(decoded twice,%252eis.) and%00. A;in the query string is fine. - Numeric hosts in one form only. Browsers and libraries read
2130706433,0177.0.0.1,0x7f.1and127.1as127.0.0.1. A host whose last label is numeric must be a canonical dotted-decimal IPv4 address, so no two parsers can disagree about which address it names. - The allow-list, as above.
- Resolve once, check every answer. The host name is resolved (A and AAAA). If any
answer is loopback, private (RFC 1918), link-local (
169.254.0.0/16, which holds the metadata address169.254.169.254, andfe80::/10), CGNAT (100.64.0.0/10), unspecified, multicast, broadcast, reserved, documentation, benchmarking, unique-local (fc00::/7, which holdsfd00:ec2::254), or an IPv6 address embedding such an IPv4 address (::ffff:a.b.c.d, NAT6464:ff9b::/96, 6to42002::/16, Teredo), the request is refused. The error does not say which address it was. - Connect to the address that was checked. The socket is opened to the vetted IP; the TLS
server name, the certificate check and the
Hostheader use the host name. A DNS server that answers with a public address for the check and a private one for the connection (DNS rebinding) gets nowhere: there is no second lookup. - No proxy, no guest headers.
http.clientdoes not readHTTP_PROXY/HTTPS_PROXY, so the environment cannot route requests elsewhere. The guest cannot pass headers or options (the tool takes exactly one argument); the request carriesHost,User-Agent,Accept-Encoding: identity,Connection: closeand the host's fixed headers. - Bounded response. The body is read in chunks up to
max_response_bytes, whateverContent-Lengthsays (a server claiming a small length cannot make it read more; a server claiming a huge one cannot make it allocate). No transparent decompression, so no zip bombs. The status line and headers are bounded byhttp.client(64 KiB per line, 100 headers). - One deadline.
timeoutcovers resolution, connection, TLS, every hop and the body. A watchdog closes the socket at the deadline, so a server dripping one byte at a time cannot hold the call open. - Redirects are never followed by the HTTP library. Each
Locationgoes back through steps 1 to 8: at mostmax_redirectshops, never fromhttpstohttp, by default only to the same scheme, host and port, and the host's fixed headers are not sent to another origin (withredirects="allow-list").
What this does not cover
- What the allowed hosts do. An allow-listed API is reachable with the host's credentials; the guest can call any GET endpoint under the allowed prefix as often as the budget allows. Scope entries tightly, and use separate tools (with separate headers) for hosts with different credentials.
- An allowed host that is itself an open redirect or a proxy. With the default
redirects="same-origin"an open redirect on the allowed host cannot leave it, but an endpoint that fetches a URL server-side is out of reach of any client-side check. - Exfiltration through the URL. The guest chooses the path and the query string, so it can send what it knows to an allowed host. If that matters, allow only hosts you trust with the session's data.
- Request methods other than GET, request bodies and cookies: not supported.
- Resolver slowness. The system resolver cannot be cancelled. Lookups run on a pool shared by
every
HttpFetchin the process (pydeno.tools.http_fetch.DNS_THREADS, 8 by default, read when the first lookup starts). A lookup that runs past the deadline raisesHttpFetchTimeoutin its caller, but keeps its thread until the OS resolver gives up, often 10 to 30 seconds. So an allow-listed name whose DNS is broken or hangs can occupy the pool and delay everyHttpFetchin the process by up to the resolver timeout; each caller still getsHttpFetchTimeoutat its own deadline, never a hang. Lookups in flight (running or waiting for a thread) are capped per process (pydeno.tools.http_fetch.DNS_MAX_PENDING, 64 by default, read withDNS_THREADSwhen the first lookup starts). A lookup still waiting for a thread is cancelled when its caller times out; a running one counts until it finishes. Past the cap a call fails at once withHttpFetchFailedinstead of queueing, so stalled lookups cannot pile up. Only allow-listed host names are ever resolved (the allow-list check comes first), so the guest cannot pick arbitrary names to stall the pool with. Passresolver=if you need a resolver with its own timeout. - TLS policy is Python's default context (system trust store, certificate and host name
verification, TLS 1.2 minimum). Pass
ssl_context=to change it. The DNS-rebinding defence has two halves: the socket goes to the vetted IP, and the certificate proves that the server there is the allow-listed host. A context withcheck_hostname=Falseorverify_mode=CERT_NONEkeeps the first and loses the second: the connection still goes only to a vetted public address, but nothing checks which server answers there. - Plain
http(when you add it toschemes) is readable and modifiable by anyone on the path.
The checks are tested against a local server (tests/test_http_fetch.py): redirects to another
host and to loopback, the 5-hop cap, a resolver that changes its answer between calls, userinfo,
decimal, octal and hex IPv4 forms, IPv6 loopback and v4-mapped addresses, the metadata address, a
body larger than the cap with and without a lying Content-Length, the deadline, non-https
schemes, CR/LF in the URL, and guest-supplied headers.