Production Checklist

kraken runs as a single container with no database, which makes it easy to start and easy to expose by accident. Work through this list before anything outside your own network can reach it.

Terminate TLS at a reverse proxy

kraken speaks plain ws:// and http:// on WS_PORT (8080) and has no TLS settings of its own. Tokens travel in the first WebSocket frame, so put kraken behind a reverse proxy or load balancer that terminates TLS and passes WebSocket upgrades through, and connect clients with wss://.

With nginx, inside the server block that holds your certificate:

location = /ws {
    proxy_pass http://127.0.0.1:8080;
    proxy_http_version 1.1;
    proxy_set_header Upgrade $http_upgrade;
    proxy_set_header Connection "upgrade";
    proxy_set_header Host $host;
    proxy_read_timeout 120s;
}

location = /health {
    proxy_pass http://127.0.0.1:8080;
}
  • Forward only /ws, and /health if your load balancer needs it. kraken also serves an internal, unsupported endpoint under /internal/ on the same port; keep it unreachable.
  • Make the proxy's idle timeout longer than the SDKs' 30 second heartbeat. kraken itself closes a connection that sends nothing for 60 seconds.
  • Any node in a cluster can take any connection, so the load balancer needs no session affinity.

Expose as little as possible

PortWho needs it
8080 (WS_PORT)Your reverse proxy only
1883 (MQTT_PORT)Nobody. The MQTT listener accepts no connections in v0.9.0.
4369 and 9100 to 9200Other kraken nodes only, when clustered. Anyone who can reach them with the cookie controls the node.
Your auth service, or @nolag/core's hostkraken only, plus whatever administers it

Change the defaults

  • Tokens. Do not ship the repository's examples/auth.json: its tokens are public. See Static Auth File.
  • AUTH_ALLOW_ALL must be false (the default).
  • ERLANG_COOKIE defaults to kraken_dev_cookie. Set a long random value, the same on every node. kraken prints the cookie in its startup log, so treat the logs as sensitive.
  • BACKEND_SECRET. Set it when using the HTTP auth or control backends, and make your service check it.
  • INTERNAL_SECRET defaults to change_me. Change it, even though the endpoint it guards is not for use.
  • RECORD_MESSAGES is true in the image. Nothing can read recorded messages back in v0.9.0, so unless you have a use for message ids and acknowledgements, set it to false and save the memory.
  • With @nolag/core: keep SIGNING_KEY_ENCRYPTION_KEY safe and backed up (losing it makes every signing key unusable), and never expose the example host as it is: it authenticates nobody. See Full Stack.

Health checks

GET /health answers 200 with {"status":"ok"} once the HTTP listener is up. It does not check the auth backend, the MQTT broker or the cluster. kraken's Docker image runs it as its HEALTHCHECK every 30 seconds.

@nolag/core's example host answers GET /health with {"status":"ok","database":"up"}, or degraded when Postgres is unreachable.

Logs

kraken logs to standard output at a single level, with no setting to quieten it in v0.9.0. It logs every connection and disconnection, actor ids, topic addresses, and refused subscribes together with the actor's granted patterns. When its own output falls far behind, it drops log lines rather than slowing the broker down.

Resources

  • Each WebSocket connection is an Erlang process, and the VM is started with a limit of a million processes.
  • Each connection holds a socket, so raise the container's open-file limit for large numbers of connections.
  • Memory grows with connections, with retained messages (the last one per topic, held for an hour by the syn broker), and with recorded messages when RECORD_MESSAGES is on (up to STORE_MAX_MESSAGES, 10,000 by default, for STORE_TTL_SECONDS, an hour by default).
  • We have not published capacity figures for v0.9.0. Load test with your own traffic shape: the syn broker checks every subscription group for wildcard matches on each publish, so many distinct subscriptions cost more than many messages.

Restarts and upgrades

Restarting a node drops all of its connections. The SDKs reconnect by themselves (the JavaScript and Python SDKs after about five seconds by default). What they get back is covered below: in short, subscribe in your connect handler, and expect messages sent during the gap to be lost.

What kraken does not do

Things a hosted or commercial broker might do for you, and kraken v0.9.0 does not. The ones marked measured we checked against a running v0.9.0.

  • No per-token limits. Every connection may publish 50 messages a second, and every payload may be up to 921,600 bytes. Per-token values (rateLimit and maxMessageSizeBytes in auth.json, max_message_size_bytes from an auth service) and MAX_MESSAGE_SIZE are ignored. Measured.
  • No message history or replay. A client that connects or subscribes late receives only new messages, plus the last retained message on a topic if one was published with retain. There is no API to read earlier messages, even when recording is on.
  • No offline queue on the default broker. A message published while a subscriber is disconnected never reaches it. Measured. Only the MQTT broker backend can keep a session for a disconnected client, and only for actors whose auth backend asks for one.
  • No automatic restore of subscriptions, unless your auth backend supplies them. After a reconnect kraken restores the subscriptions the auth backend lists for the actor, and nothing else. The repository's demo tokens and the nolag-core quickstart list none, so JavaScript and Python clients reconnected and then received nothing until they subscribed again. Measured. The Go SDK subscribes again by itself. @nolag/chat, run against the static file, re-joined its online lobby after a broker restart but stopped receiving room messages. Measured. Subscribe in your client's connect handler; see the Quick Start.
  • No QoS on the default broker, and no end-to-end delivery guarantee on any. syn ignores QoS. With the MQTT backend, QoS applies to the hop between kraken and the broker. The only acknowledgement a publisher gets is that kraken accepted the message. See Quality of Service.
  • No echo suppression by default. A connection subscribed to a topic receives its own publishes unless it publishes with echo: false. Measured.
  • No instant revocation. Live connections keep their grants until their next revalidation, up to about ten minutes after a token is revoked, and new connections may be accepted from the 30 second auth cache. Measured for revalidation.
  • No TLS. See above.
  • No MQTT device ingress. The MQTT listener on port 1883 accepts no connections in v0.9.0. Measured.
  • No gossip discovery. CLUSTER_STRATEGY=gossip does not form a cluster; use epmd or dns. Measured.
  • No persistent presence or wake-up. Presence lasts exactly as long as the connection. The presence store and wake plugin slots default to doing nothing, and the in-memory and HTTP modules kraken includes for them are marked in the source as for development and testing only.
  • No signed webhooks. Hydration and trigger webhook requests carry only the headers you configure, so authenticate them with a secret header. A webhook call that fails with a server error or a connection failure is attempted up to three times in all; a 4xx answer is not retried.

Next steps