Steps to reproduce
-
Create a gateway with dstack 0.22.2. Deploy one authenticated HTTP service behind that gateway. It must return 200 for GET /. The reproduced configuration used the default Python soft FD limit of 1024.
-
Set the public service URL routed through that gateway, as printed by dstack apply, and a token with access to its project. Replace the example address with your service's actual address:
export SERVICE_URL="https://my-service.example.com/"
export PROJECT_TOKEN="<your-project-token>"
Confirm GET / returns 200 for this service URL using that token. The load commands below require hey and Python 3.11+.
-
Send valid-token load, then wait 70 seconds:
ulimit -n 8192 # load generator only
hey -z 3m -c 1024 -q 0.2 -t 20 -disable-keepalive \
-H "Authorization: Bearer $PROJECT_TOKEN" "$SERVICE_URL"
sleep 70
-
Send requests with a different invalid token every time to bypass the authorization cache: 1,024 workers, a 5-second period, a 20-second timeout, for 120 seconds. These requests normally return 403.
Run the uncached load
python3 - <<'PY'
import asyncio, collections, contextlib, os, ssl, time, uuid
from urllib.parse import urlsplit
u = urlsplit(os.environ["SERVICE_URL"]); counts = collections.Counter(); run_id = uuid.uuid4().hex
host, port = u.hostname, u.port or (443 if u.scheme == "https" else 80)
tls = ssl.create_default_context() if u.scheme == "https" else None
path = (u.path or "/") + ("?" + u.query if u.query else "")
async def worker(i, start, end):
n, next_start = 0, start
while time.monotonic() < end:
writer = None
try:
async with asyncio.timeout(20):
reader, writer = await asyncio.open_connection(host, port, ssl=tls)
writer.write((f"GET {path} HTTP/1.1\r\nHost: {u.netloc}\r\n"
f"Authorization: Bearer invalid-{run_id}-{i}-{n}\r\nConnection: close\r\n\r\n").encode())
await writer.drain(); line = await reader.readline(); await reader.read()
counts[line.split()[1].decode() if line else "closed"] += 1
except Exception as exc: counts[type(exc).__name__] += 1
finally:
if writer:
writer.close()
with contextlib.suppress(Exception): await writer.wait_closed()
n += 1; next_start = max(next_start + 5, time.monotonic())
await asyncio.sleep(max(0, min(next_start, end) - time.monotonic()))
async def main():
start = time.monotonic(); await asyncio.gather(*(worker(i, start, start + 120) for i in range(1024)))
asyncio.run(main()); print(dict(counts))
PY
-
After the generator exits, wait 70 seconds and send a normal request:
sleep 70
curl -i -m 30 -H "Authorization: Bearer $PROJECT_TOKEN" "$SERVICE_URL"
This sequence reproduced the persistent failure in an isolated test. Gateway capacity and timing affect whether it triggers.
Expected behaviour
Valid requests succeed again after load stops, without restarting the gateway.
Actual behaviour
Services behind the gateway keep returning nginx 500 after the load stops. The dstack CLI and server API remain healthy. Restarting the gateway restores service.
The service used for the load test still returned 500 after the 70-second wait, with authorization PoolTimeout errors in the gateway logs.
Suggested fix
- Set
LimitNOFILE=65535 in dstack.gateway.service (nginx's limit does not apply to the Python process).
- In
gateway/services/server_client.py, close auth connections left marked busy after their request has timed out or been cancelled, so later requests can proceed.
- Build the list of available auth connections once when processing waiting requests, instead of rebuilding it for every request.
dstack version
0.22.2; HTTPX 0.28.1, HTTPcore 1.0.5, AnyIO 4.3.0.
Steps to reproduce
Create a gateway with dstack 0.22.2. Deploy one authenticated HTTP service behind that gateway. It must return 200 for
GET /. The reproduced configuration used the default Python soft FD limit of 1024.Set the public service URL routed through that gateway, as printed by
dstack apply, and a token with access to its project. Replace the example address with your service's actual address:Confirm
GET /returns 200 for this service URL using that token. The load commands below requireheyand Python 3.11+.Send valid-token load, then wait 70 seconds:
Send requests with a different invalid token every time to bypass the authorization cache: 1,024 workers, a 5-second period, a 20-second timeout, for 120 seconds. These requests normally return 403.
Run the uncached load
After the generator exits, wait 70 seconds and send a normal request:
This sequence reproduced the persistent failure in an isolated test. Gateway capacity and timing affect whether it triggers.
Expected behaviour
Valid requests succeed again after load stops, without restarting the gateway.
Actual behaviour
Services behind the gateway keep returning nginx 500 after the load stops. The dstack CLI and server API remain healthy. Restarting the gateway restores service.
The service used for the load test still returned 500 after the 70-second wait, with authorization
PoolTimeouterrors in the gateway logs.Suggested fix
LimitNOFILE=65535indstack.gateway.service(nginx's limit does not apply to the Python process).gateway/services/server_client.py, close auth connections left marked busy after their request has timed out or been cancelled, so later requests can proceed.dstack version
0.22.2; HTTPX 0.28.1, HTTPcore 1.0.5, AnyIO 4.3.0.