Evaluate deployment hardening
beryl disables rate and connection limits by default because each deployment needs different values. These defaults are suitable for development. A hostile or faulty client can degrade a server that has no traffic controls. beryl logs a startup warning when all controls are off. This guide describes defensive controls and testable starting points; it does not provide production-validated capacity recommendations.
Controls enabled by default
Section titled “Controls enabled by default”Even with no configuration, beryl enforces:
- Local work admission: finite router, socket, worker, and presence budgets, plus callback-result limits. See overload handling for defaults, Result APIs, and memory exclusions.
- Post-receipt frame size: once the transport has assembled a complete
inbound WebSocket frame, beryl closes the connection if it exceeds 1 MiB
(
with_max_inbound_frame_bytesto adjust). - Topic and event lengths: topics over 256 bytes and event names over 64
bytes are rejected before reaching your app
(
with_max_topic_length,with_max_event_length). - Joined-topic cap: a socket may join at most 1000 topics
(
with_max_joined_topics_per_socket). - Protocol checks: reserved
phx_*events, reservedberyl:*topics, and messages carrying a stalejoin_refare dropped. - Heartbeat eviction: sockets that stop sending heartbeats are evicted
and their connections closed (60 s window by default,
with_heartbeat). - Outbound queue budget: each connection may reserve at most 256 frames
and 1 MiB of payload data while writes are pending. beryl closes a slow
connection instead of dropping frames. Use
with_outbound_limitson the transport config to adjust both limits.
The frame-size check limits decoding and routing work. It does not limit transport memory because buffering and reassembly occur before beryl receives the frame. At deployment, set a WebSocket frame or message limit in the reverse proxy or load balancer. Set it at or below beryl's limit. Also set a matching HTTP request or body limit for the upgrade.
Choose initial limits
Section titled “Choose initial limits”let config = beryl.config(wire.phoenix_codec()) // Drop over-rate complete frames before decoding them. |> beryl.with_frame_rate(per_second: 150, burst: 300) // Example only. Measure your busiest legitimate client and tune this value. |> beryl.with_message_rate(per_second: 100, burst: 200) // Example only. Tune this for your reconnect and rejoin behavior. |> beryl.with_join_rate(per_second: 10, burst: 20) // Cap connection attempts per client IP. This allowance survives // disconnects and app runtime restarts. |> beryl.with_connection_rate_per_ip(per_second: 5, burst: 10) // Concurrent connections per client IP. Size to your expected // clients-behind-one-NAT worst case; see the caveat below. |> beryl.with_max_connections_per_ip(max_connections: 100) // Per-system ceiling on this node across all IPs. Size all systems on a // node to its combined process/socket/runtime budget; see below. |> beryl.with_max_connections(max_connections: 10_000)
let assert Ok(websocket_config) = server.default_config("/socket/websocket") |> server.with_outbound_limits(max_frames: 256, max_bytes: 1_048_576)with_frame_rate and with_message_rate are independent. Joins use frame and
join quota, but not message quota. Leaves and heartbeats use frame and message
quota. Configure both limits. Give the frame limit more capacity for protocol
traffic and malformed frames. If a limiter drops a heartbeat, the runtime does
not refresh the heartbeat deadline. Continued excess traffic then causes
heartbeat eviction.
Size both rate and burst allowances with enough headroom for legitimate client traffic plus heartbeats. A limit that admits an application's normal events but leaves no heartbeat capacity can evict healthy clients during ordinary bursts.
The outbound budget is enforced before a frame enters the connection mailbox. For both text and binary codec output, it charges the referenced BEAM binary size, not only the logical frame length. For example, a 128-byte slice of an 8 MiB binary uses 8 MiB of the budget. Each frame is charged separately, even if frames share a backing binary. beryl does not copy these slices.
Ok from the runtime send callback means that the frame was
enqueued; it does not mean that the transport wrote it or that the peer
received it. Releasing a runtime work reservation does not release the
transport reservation. Successful writes release their frame and byte
reservations. Write errors and connection close release all remaining
reservations. If a connection exceeds either limit, beryl closes it so the client can reconnect
and resynchronize instead of receiving a stream with silently dropped frames.
This is not a total-memory cap. Process heaps, WebSocket framing, and transport or kernel buffers are outside this budget.
Set the byte budget above the largest legitimate encoded frame's referenced binary size, or copy small slices in your codec to avoid retaining large backing binaries. Size the frame budget for short write stalls, not sustained client outages. The current defaults allow 256 frames and 1 MiB per connection. Validate whether those defaults fit your workload.
Optionally, with_channel_rate adds a per-socket-per-topic limit on top of
the global per-socket message rate, useful when a single busy topic must not
starve others.
Combine per-IP rate and connection limits
Section titled “Combine per-IP rate and connection limits”with_connection_rate_per_ip caps how quickly each peer can open connections,
which limits how often reconnects can refresh per-connection frame and message
bursts. The connection limiter checkpoints these per-IP buckets in an ETS
table with an heir. They survive disconnects, router restarts, and
limiter-worker restarts. They do not survive shutdown or replacement of the
enclosing beryl supervisor, or a node restart. Idle buckets expire after their
allowance has fully refilled.
with_max_connections_per_ip separately throttles a single peer's concurrent
connections, while with_max_connections caps concurrent connections across
all IP addresses for one beryl system on one BEAM node. A connection must pass
every configured rate and concurrency limit; otherwise the transport rejects it
with 429 before allocating any long-lived socket or runtime state. Freed
concurrency capacity is reclaimed on normal close, transport failure, heartbeat
eviction, crash, and setup failure.
A per-IP limit cannot stop many source addresses. A botnet or a host that rotates IPv6 addresses can open a few connections from each address. Together, these connections can exhaust the system's capacity. The per-system ceiling limits the total number of connections across all IP addresses.
Each independently constructed Sockets system has its own limiter. Two
systems on one node can therefore each admit up to their configured limit. With
one system per node, a load-balanced cluster of N nodes has an effective
ceiling of roughly max_connections × N. If a node runs multiple systems, size
their combined limits against that node's capacity. Use your load balancer's
own global connection/rate controls when you need a cluster-wide cap.
Limits behind proxies and shared IP addresses
Section titled “Limits behind proxies and shared IP addresses”The connection limit uses the TCP peer address. It ignores forwarded
headers such as X-Forwarded-For because clients can forge them. This has two
effects:
- Behind a reverse proxy or load balancer, every connection appears to come from the proxy's IP, so a per-IP cap would throttle all clients together. Enforce per-client limits at the proxy layer instead.
- Users behind carrier-grade NAT or a corporate gateway share one IP. Set the cap high enough for your worst legitimate case, or leave it off and rely on rate limits.
Reconnects reset per-connection limits
Section titled “Reconnects reset per-connection limits”Transports store frame-rate buckets per connection; socket actors store
message and join buckets. When those owners exit, their buckets disappear.
A client that reconnects gets a fresh allowance. These limits bound the damage
of a single connection; they do not preserve a client's quota across reconnects.
Configure with_connection_rate_per_ip to limit repeated reconnects from one
peer IP. Keep infrastructure-level controls (load balancer connection/request
limits, WAF rules) for limits that must survive beryl restarts and for attackers
rotating source addresses.
Check origins and authenticate
Section titled “Check origins and authenticate”server.default_configuses same-origin checks by default. Usewith_allowed_originswhen you need an explicit list of browser origins; unexpected origins are rejected before the WebSocket handshake.with_on_connectauthenticates the connection once, before upgrade. Reject unauthenticated clients with a 403 rather than at join time.- Authorize each topic in your update's
Joinarm; clients cannot send events to topics they have not joined.
Secure the Erlang cluster
Section titled “Secure the Erlang cluster”Before enabling Erlang distribution, restrict its ports to trusted hosts and configure mutually verified TLS. Apply these controls even to a single node, in development and staging as well as production, whether or not it runs beryl.
beryl PubSub and presence replication use Erlang distribution. Trust every connected peer. Erlang peers can run arbitrary code on connected nodes. A hostile peer can compromise the full cluster. Topic access, broadcasts, internal traffic, and presence state are only part of that access. Channel authorization and WebSocket controls do not apply to distribution messages. They protect only WebSocket clients.
Trust for client and cluster traffic
Section titled “Trust for client and cluster traffic”| Source | Trust level | Protected by |
|---|---|---|
| WebSocket clients | Untrusted | with_on_connect authentication, Join authorization, and Message handling in update |
| Erlang distribution peers | Fully trusted | Network isolation + mutually verified TLS distribution (cookies prevent accidental cross-cluster connections only) |
The Erlang cookie does not secure distribution
Section titled “The Erlang cookie does not secure distribution”Set a long, randomly generated cookie in vm.args or the
RELEASE_COOKIE environment variable. Erlang
uses the cookie to distinguish clusters
and prevent accidental cross-cluster connections; its handshake is not
cryptographically secure authentication against an adversary. A strong cookie
prevents accidental matches and makes offline guessing harder, but it must
never be the security control. Keep distribution ports isolated and use
mutually verified TLS distribution.
Generate a strong cookie with, for example:
openssl rand -base64 48Use mutually verified TLS distribution
Section titled “Use mutually verified TLS distribution”Use TLS distribution with mutual certificate verification when enabling distribution, including on private networks. See the Erlang TLS Distribution guide for setup instructions.
Restrict EPMD and distribution ports
Section titled “Restrict EPMD and distribution ports”The EPMD port (default 4369) and the Erlang distribution listen range must not be reachable from untrusted networks. Use firewall rules or security groups to allow only cluster-internal traffic. You can fix the distribution port to a known value for simpler rules:
# vm.args-kernel inet_dist_listen_min 9100-kernel inet_dist_listen_max 9100Or set ERL_DIST_PORT when using Elixir/Mix releases. Restrict both 4369
and your chosen distribution port at the network layer.
Keep untrusted nodes outside the cluster
Section titled “Keep untrusted nodes outside the cluster”Adding a node to a cluster lets it run arbitrary code on connected nodes. It
can also access all pg groups and presence state. Do not connect beryl to
nodes you do not trust.
See SECURITY.md for the full distribution security reference.
Runtime capacity
Section titled “Runtime capacity”- The runtime uses one router actor per channel system and one actor per connected socket. The transport drops oversized frames and, when configured, over-rate traffic before routing. Applications that send one event to many subscribers may need several channel systems divided by topic so one router does not handle every subscriber.
- The runtime always starts supervised (
child_spechas no unsupervised mode); restarts are bounded at 3 per 5 seconds, after which the crash propagates through the application's supervision tree. See the Supervision guide. - Rate limits do not bound worker input, pending action reports, router fan-out, or outbound connection queues. See Queue limits and overload for current controls and non-guarantees.
