Amazon S3 Tables now accept all data types defined in Apache Iceberg version 3, and they expose the V3 capabilities of deletion vectors, row lineage, and the extended type set (variant, nanosecond timestamps, geometry, geography, unknown). This change lets you create new V3‑format tables or upgrade existing V2 tables without leaving the S3 Tables service, while the managed compaction and maintenance pipelines continue to run automatically.
Key technical changes
The update adds the following Iceberg V3 features to S3 Tables:
- Deletion vectors replace positional delete files with a compact binary format, reducing the number of small files generated by row‑level deletes.
- Row lineage automatically adds
_row_idand_last_updated_sequence_numbercolumns to each record, enabling downstream pipelines to identify changed rows without full scans. - Extended data types – variant, nanosecond‑precision timestamps (with time zone), geometry, geography, and unknown – are now natively supported.
All of these are available when you set the Iceberg table property format-version='3' or when you alter an existing V2 table in place.
Why the change matters to practitioners
AI and analytics engineers often ingest semi‑structured JSON events, geospatial logs, or high‑resolution timestamps. Previously they had to store such values as strings or integers, incurring extra storage, parsing overhead, and schema‑evolution friction. With native variant and geospatial types, the engine can shred the data into columnar form and collect statistics, which improves file pruning and reduces I/O.
Compliance or data‑retention deletions that touched millions of rows previously produced thousands of positional delete files, slowing queries until a compaction cycle ran. Deletion vectors collapse those deletes into a single binary file, cutting compaction time and query latency.
Row lineage provides a built‑in audit trail that can be queried directly, simplifying downstream change‑data‑capture pipelines and reducing the need for custom timestamp columns or external tracking services.
Operational considerations
Upgrading a V2 table to V3 is performed in place; the service continues to handle automatic compaction, replication, and Intelligent‑Tiering. However, practitioners should be aware of the following operational points:
- Enabling merge‑on‑read write modes (e.g.,
write.delete.mode='merge-on-read') is required to generate deletion vectors. - Row‑level metadata columns are added automatically; downstream code that selects
*will now see the extra fields. - Because S3 Tables manage compaction, any custom compaction schedules should be reviewed to avoid overlapping with the service’s maintenance windows.
Example workflow for a clickstream table demonstrates the new capabilities:
CREATE TABLE my_catalog.namespace.clickstream (
event_id bigint,
event_time timestamp,
user_id string,
payload variant
) USING iceberg TBLPROPERTIES ('format-version' = '3');
INSERT INTO my_catalog.namespace.clickstream VALUES
(1, current_timestamp(), 'user-42', PARSE_JSON('{"action":"purchase","amount":99.99}')),
(2, current_timestamp(), 'user-17', PARSE_JSON('{"action":"page_view","url":"/products/webcam"}'));
SELECT event_id, user_id,
variant_get(payload, '$.action', 'string') AS action,
variant_get(payload, '$.amount', 'double') AS amount
FROM my_catalog.namespace.clickstream
WHERE variant_get(payload, '$.action', 'string') = 'purchase';
ALTER TABLE my_catalog.namespace.clickstream SET TBLPROPERTIES (
'write.delete.mode' = 'merge-on-read',
'write.update.mode' = 'merge-on-read',
'write.merge.mode' = 'merge-on-read'
);
DELETE FROM my_catalog.namespace.clickstream WHERE user_id = 'user-42';
Related CloudNinjas coverage: AWS.
What This Means For Practitioners
Adopt Iceberg V3 on S3 Tables when you need native handling of semi‑structured or geospatial data, or when row‑level delete performance is a bottleneck. Plan a phased upgrade of existing V2 tables, validate that downstream queries tolerate the added lineage columns, and adjust any custom compaction or maintenance scripts to defer to the managed service. Monitor the size and count of deletion vector files after enabling merge‑on‑read writes to confirm the expected reduction in small‑file churn.

