Postgres Logical Replication Slot Mechanics
Understanding how logical replication slots track database changes through WAL positions.

Every data modification in PostgreSQL is written to the Write-Ahead Log before it ever touches the heap, the actual table data files on disk. That ordering is not incidental. It makes the WAL the authoritative, sequential record of everything that has happened to the database, and it is the reason any replication technology built on top of PostgreSQL has to start here rather than anywhere else. Each record in that log describes a physical change: which bytes moved at which page offset, tagged with a Log Sequence Number, a 64-bit pointer that identifies both the WAL segment and the exact offset within it. That LSN gives the log its order, so two servers, or two consumers, can compare their position in the stream, because every byte written has a unique, monotonically increasing address.
Physical replication ships the raw WAL bytes straight to a standby server, so the copy it produces is byte-identical to the primary. It is fast, but rigid: the whole cluster replicates as a unit, there is no way to replicate a single table, and both sides have to run the same major version of PostgreSQL. That rigidity comes from the same property that makes the WAL trustworthy. The format is physical: it describes storage-level changes to pages, not logical operations on rows. A data warehouse, a message queue, or any other external system has no way to interpret a page-level byte diff. It needs to know that a row in the orders table had its status column change from pending to shipped, not that a handful of bytes at a particular offset were overwritten. That gap between what the WAL physically contains and what an external consumer can actually use is the problem logical decoding exists to solve.
Logical Decoding: From Physical WAL Records to a Row-Level Change Stream
The PostgreSQL documentation defines logical decoding as the process of extracting all persistent changes to a database's tables into a coherent, easy-to-understand format that can be interpreted without detailed knowledge of the database's internal state. What arrives on the other end is no longer a page diff but a readable account of what changed, in the order it changed, grouped by the transaction that made it happen.
That translation happens at the row level rather than the block level, and that distinction is what makes selective replication possible at all: a consumer can ask for changes to specific tables rather than the entire cluster, and those changes can land in systems that have nothing to do with PostgreSQL's on-disk format. But the decoder can only emit what the WAL was told to record in the first place, and two settings determine that. The server parameter wal_level has to be set to logical in postgresql.conf; the default value, replica, writes enough information for physical replication and point-in-time recovery but not enough for logical decoding to function. Changing it requires a server restart, so it has to be planned for a maintenance window rather than flipped on casually.
The second setting lives at the table level: REPLICA IDENTITY. It governs which columns show up in the old-row image attached to UPDATE and DELETE events, so a consumer can know what a row looked like before it changed. By default, PostgreSQL records only the primary key columns in that before-image, which is enough for most consumers to identify which row changed but not enough to see what every column held beforehand. If you set REPLICA IDENTITY FULL, you capture every column in the before-image, but you pay for it with additional WAL volume on every single update and delete against that table, so you should choose it table by table rather than apply it mechanically across a schema. One further wrinkle belongs here as well: TOAST column values, PostgreSQL's mechanism for storing oversized field values out of line, appear in the decoded new-row image only when their value actually changed. A CDC consumer comparing before and after images has to account for that, or it will misread an unchanged large column as missing data.
Output plugins: how decoded changes get formatted for the consumer
Once the decoder has reconstructed a row-level stream of changes, that stream still has to be serialized into a format the consumer on the other end can actually parse, and that job belongs to the output plugin, the final stage of the pipeline before changes leave the server. PostgreSQL's documentation describes the output plugin as the component responsible for translating the server's internal representation into the application-specific form the slot's consumer wants. Several plugins fill that role in production today, each with a different format and a different niche. pgoutput is PostgreSQL's built-in plugin, introduced in PostgreSQL 10, and it ships on every major managed Postgres service, including AWS, GCP, and Azure, without any additional installation required. It uses an efficient binary replication message format, supports filtering changes by publication, and handles transaction boundaries natively, which makes it the right default choice for production CDC work. test_decoding produces plain text output and exists mainly to help engineers learn the replication protocol or debug a pipeline, not to run production traffic. wal2json emits JSON, and CDC tools that prefer a text-based change format over a binary one tend to reach for it. decoderbufs decoderbufs serializes changes using a binary serialization format, a choice suited to high-throughput pipelines where every byte of serialization overhead matters.
The practical case for pgoutput as a default goes beyond convenience. Because it is maintained as part of the PostgreSQL core project rather than as a separate extension, it tracks changes to the WAL record format automatically as new PostgreSQL versions ship. Third-party plugins that are not kept in lockstep with those changes carry a real risk: a major-version upgrade can silently break them if their parsing logic falls out of sync with the new WAL formats. A built-in, core-maintained plugin avoids that risk, and that is why it anchors most production logical replication setups built directly on PostgreSQL. With decoding and formatting both accounted for, the question becomes how PostgreSQL tracks where each consumer stands in that stream, and that tracking is the job of the replication slot.
What a Replication Slot Is: the Two LSN Watermarks
A replication slot is PostgreSQL's bookmark in the WAL stream. It records exactly where a given consumer has gotten to, and in doing so it prevents the server from discarding WAL segments, or vacuuming away catalog rows, that the consumer might still need. Every slot carries two LSN watermarks that together define its behavior, and understanding both is the key to understanding everything a slot does and everything that can go wrong with one.
The first is restart_lsn: it marks the earliest WAL position the slot might still need. PostgreSQL treats this as an absolute floor: it will not recycle or remove any WAL segment at or after this point, no matter how old it gets or how much disk space it consumes. The second is confirmed_flush_lsn, the position the consumer has explicitly acknowledged as durably stored on its own side. That acknowledgment is the consumer telling PostgreSQL, in effect, that everything up to this LSN has been safely written somewhere else and can be forgotten by the publisher. Only after that confirmation arrives does the slot release its hold on the corresponding WAL. The distance between confirmed_flush_lsn and the server's current WAL write position, pg_current_wal_lsn(), is a computable quantity, using the function pg_wal_lsn_diff(), and it tells an operator the exact byte count of WAL the slot is holding hostage on behalf of that one consumer.
Slot state is persisted to disk only at checkpoint, not continuously. After a crash, a slot can rewind to an earlier LSN than the consumer had actually reached and re-emit changes it already sent once. The PostgreSQL documentation is direct about the consequence: logical decoding clients are responsible for handling duplicate messages themselves, because the slot guarantees at-least-once delivery at the WAL layer, not exactly-once delivery. A slot also has no awareness of its own consumer's health or identity. Different receivers can use the same slot at different times, each picking up from wherever the previous one left off, but only one receiver can consume a given slot at any single moment. A PostgreSQL instance can host multiple independent slots simultaneously, each tracking a different consumer's progress through the same stream, but that number is not unbounded: max_replication_slots caps how many slots can exist (ten, by default), and max_wal_senders caps how many simultaneous WAL-streaming connections the server will support. You need to size both deliberately before your team starts adding CDC consumers, so you don't discover them as limits after the fact.
The Publication and Subscription Model
Publications and subscriptions form the declarative layer PostgreSQL builds on top of the slot machinery just described. A slot answers one question: how far has this consumer gotten. Publications and subscriptions answer two different ones: what changes should be sent, and where should they go. A publication is defined on the publisher side and names which tables, or whether all tables, participate in replication. Because a single slot streams changes from one database as a whole, publications are what let that single stream be narrowed down to a meaningful subset of tables rather than forcing every consumer to ingest the entire database.
A subscription lives on the opposite side, on the subscriber, and it holds the connection details back to the publisher along with the name of the publication it wants to follow. By default, creating a subscription automatically creates a replication slot on the publisher, named after the subscription itself, which is the mechanism that ties the declarative layer back down to the LSN-tracking primitive underneath it. The command CREATE SUBSCRIPTION performs several steps in a fixed order: it creates that replication slot first, then runs an initial COPY of every table in the publication to establish a starting baseline, and only after that copy begins streaming the WAL changes that arrived once the copy started. That ordering is deliberate. If you create the slot before the copy, you lose no change that occurs during the copy window, in the gap between the two steps.
For large tables, the default behavior ties a synchronous copy to subscription creation, and that becomes a liability instead of a convenience. If you have large tables, you are generally better off setting create_slot = false and calling pg_create_logical_replication_slot manually, so you can stage the initial data copy separately, outside the critical path of subscription setup. External CDC tools that read from PostgreSQL and write into data warehouses typically connect directly to a replication slot using pgoutput or another output plugin, bypassing the publication and subscription apparatus. The underlying slot mechanism is identical either way. Only the consumer attached to it differs.
Why an idle or abandoned slot is a silent disk-filling hazard
Everything a slot is built to guarantee, durable delivery, crash safety, protection against premature data loss, rests on one behavior: PostgreSQL will not delete a WAL segment that a replication slot still references. That guarantee is also where the danger lives. A slot that is disconnected, forgotten, or simply never cleaned up after a test causes WAL to accumulate on disk without any built-in limit, quietly, until the primary server runs out of space and stops accepting writes.
As long as restart_lsn fails to advance, every WAL segment behind that point is locked in place and cannot be recycled, no matter how much disk pressure builds. A subscriber that goes offline permanently, without anyone dropping its slot, leaves that slot pinning every WAL segment generated since the last moment it confirmed receipt. This is, by a wide margin, the most common operational failure mode teams run into with PostgreSQL logical replication: a slot gets created to test a pipeline and is never removed, or a CDC consumer crashes and nobody restarts it, and the slot keeps holding WAL in its place regardless.
The damage is not confined to WAL storage. A slot also tracks the oldest transaction still visible to its consumer, and that tracking prevents VACUUM from reclaiming dead tuples newer than that transaction. A stalled consumer therefore inflates both pg_wal disk usage and ordinary table bloat at the same time, because the same watermark that protects WAL segments also blocks routine cleanup. In the most extreme cases, unchecked slot retention can force PostgreSQL to shut itself down entirely to prevent transaction ID wraparound, a full stop rather than a gradual slowdown. None of this trips an alarm on its own. No default alert fires when a slot's lag grows, so the disk fills gradually and then, past a threshold, all at once. Dropping a subscription on the subscriber side compounds the problem if the publisher connection is already dead: the slot on the publisher is not automatically removed, and it continues holding WAL until someone explicitly drops it there.
The operational controls for keeping slot lag from becoming a production incident
PostgreSQL gives you a layered set of tools for keeping slot lag inside safe bounds, moving from prevention through detection to recovery, so if you want to run slots safely in production, you need all three layers together, not just one.
The first layer caps how much WAL a single slot is allowed to hold. Since PostgreSQL 13, you can use the parameter max_slot_wal_keep_size to set a ceiling on accumulated WAL per slot. Once a slot's retained WAL crosses that ceiling, PostgreSQL invalidates the slot so the disk does not fill without bound. For most production systems, that tradeoff is the right one: a consumer that reconnects to an invalidated slot has to re-snapshot its data from scratch, which carries a real cost, but the database itself stays online and keeps accepting writes. The invalidation itself is explicit rather than silent. The consumer gets a clear error telling it that the slot no longer holds the WAL it needs, instead of discovering corrupted or missing data later.
The second layer is active monitoring rather than a hard cutoff. Querying pg_stat_replication shows write_lsn, flush_lsn, and replay_lsn for active replication connections, revealing how far behind each one currently sits. For logical slots specifically, comparing confirmed_flush_lsn from pg_replication_slots against the server's current position, pg_current_wal_lsn(), and taking the difference with pg_wal_lsn_diff(), produces the exact byte volume of WAL still waiting to be consumed. That figure belongs on a dashboard with an alert attached to it, since lag that climbs steadily over time is the clearest available signal that a consumer has fallen behind or stopped.
The third consideration concerns databases with uneven or sporadic write activity. If a database goes long stretches without a write, its slot's LSN will not advance during that time, whether or not the consumer is healthy, so a simple lag metric alone can misrepresent what is actually happening. Distinguishing a genuinely stalled consumer from a database that is merely quiet requires pairing LSN lag with other signals, such as how recently the consumer last connected, rather than treating a static or slow-moving LSN figure as proof of trouble on its own. Used together, a hard retention ceiling, continuous lag monitoring, and an awareness of normal write patterns turn slot management from a source of silent risk into a routine, observable part of running PostgreSQL in production.
Sources
- PostgreSQL: Documentation: 18: 47.2. Logical Decoding Concepts
- PostgreSQL: Documentation: 18: 19.6. Replication
- Postgres Write-Ahead Logs
- PostgreSQL: Documentation: 18: 19.5. Write Ahead Log
- PostgreSQL: Documentation: 18: 53.20. pg_replication_slots
- PostgreSQL: Documentation: 18: 29.9. Architecture
- The wal2json plugin - Neon Docs
- Postgres Replication Slot 101: How to Capture CDC Without Breaking Production


