Pg encoding context

From PostgreSQL wiki
Jump to navigationJump to search

NB! THIS IS WORK IN PROGRESS

High Level Design: PostgreSQL Encoding Context (pg_encoding_context)

1. Executive Summary & Problem Statement

PostgreSQL currently supports built-in TOAST compression using algorithm identifiers stored in pg_attribute.attcompression ('p' for PGLZ, 'l' for LZ4). However, modern data storage, AI vector embeddings, columnar formats, and secure enterprise workloads require a generalized mechanism for context-aware and stateful encoding and decoding:

  1. Pre-Trained Compression Dictionaries: Modern algorithms like Zstandard and LZ4 achieve significant compression ratios on short-to-medium records (such as JSON documents or catalog rows) only when provided with pre-trained dictionaries.
  2. Asymmetric Codecs: Algorithms like Huffman coding, Finite State Entropy (FSE), and prefix trees require distinct representations for encoding (symbol-to-bit frequency tables) and decoding (bit-to-symbol lookup trees).
  3. Beyond Pure Compression: "Compression" is a subset of "Encoding". The same (context, input) -> output paradigm applies directly to:
  - Schema-driven binary serialization (Apache Avro, Protocol Buffers, FlatBuffers).
  - Column-level asymmetric / envelope encryption (write-only public-key ingest).
  - Vector quantization (Product Quantization for high-dimensional embeddings in pgvector).
  - NLP/LLM tokenization (BPE, WordPiece vocabularies).
  - Categorical dictionary interning (deduplicating repetitive strings into compact integer IDs).

This design introduces pg_encoding_context: a first-class, cataloged PostgreSQL object that encapsulates encoding/decoding functions, configuration parameters, and dual asymmetric tables (enctab and dectab).


2. Architecture & Design Principles

flowchart TD
    subgraph DDL ["DDL & Catalogs"]
        CCTX["CREATE ENCODING CONTEXT<br/>(pg_encoding_context)"]
        ATTR["ALTER TABLE ... ALTER COLUMN<br/>SET ENCODING CONTEXT<br/>(pg_attribute.attencoding_context)"]
        CCTX -->|referenced by OID| ATTR
    end

    subgraph Runtime ["Backend Runtime"]
        CACHE["Backend-Local Cache<br/>(HTAB in TopMemoryContext)<br/>keyed by context_id"]
        CCTX -.->|cached & invalidated| CACHE
        EXEC["Heap / TOAST Write Pipeline<br/>(toast_save_datum)"]
        DETOAST["Direct TOAST Read Pipeline<br/>(detoast_attr / slice)"]
    end

    subgraph Storage ["Direct TOAST Storage"]
        CHUNK["Direct TOAST Table<br/>(chunk_id, chunk_seq, chunk_data,<br/>chunk_tids, chunk_tid_offsets, <b>chunk_ectx_id</b>)"]
    end

    ATTR -->|fetch active context| EXEC
    EXEC -->|compress using ectx| CHUNK
    CHUNK -->|self-describing chunk_ectx_id| DETOAST
    DETOAST -->|lookup context| CACHE

Core Design Decisions

  1. Decoupled 2-Patch Architecture:
  - Patch 1: Generic pg_encoding_context catalog, C API, backend-local cache, and SQL DDL.
  - Patch 2: Direct TOAST integration with slice boundary context checkpoints (chunk_ectx_id).
  1. Option B Calling Convention (INTERNAL Pointer):
  - Handler functions pass EncodingContext directly as an INTERNAL pointer for zero-copy C execution:
    - encfunc(internal context, anyelement datum) RETURNS bytea
    - decfunc(internal context, bytea encoded_data) RETURNS anyelement
  1. Dual Asymmetric Tables (enctab vs. dectab):
  - Distinct payloads for encoding (symbol-to-code tree, writer schema, public key, quantization codebook) and decoding (decoding tree, reader schema, private key, reconstruction table).
  - Symmetric algorithms leave dectab as NULL and automatically fall back to enctab.
  1. Slice-Agnostic Decompression:
  - The decoding function is slice-agnostic; it simply decodes the byte buffer passed to it.
  - Slicing is solved at the Direct TOAST physical layer by chunking data at slice boundaries and tagging each chunk with chunk_ectx_id.
  1. Write vs. Read Path Independence:
  - Write: Governed by the column's current configuration (pg_attribute.attencoding_context).
  - Read: Self-describing from the stored Direct TOAST chunk (chunk_ectx_id), making reads resilient against table alterations or retrained contexts.

3. System Catalog: pg_encoding_context

3.1 Catalog Layout

CATALOG(pg_encoding_type,9441,EncodingTypeRelationId)
{
    Oid         oid;            /* oid */
    NameData    enctypname;     /* type name */
    Oid         enctypnamespace;/* schema namespace */
    Oid         enctypowner;    /* owner role */
    char        enctypkind;     /* high-level kind: 'c', 's', 'e', 'q', 't', 'd', 'm' */
    NameData    enctyptype;     /* specific algorithm: "zstd", "avro", "huffman", etc. */
    regproc     enctypencfunc;  /* optional custom encode function OID */
    regproc     enctypdecfunc;  /* optional custom decode function OID */
} Form_pg_encoding_type;

CATALOG(pg_encoding_context,9431,EncodingContextRelationId)
{
    Oid         oid;            /* oid */
    NameData    ectxname;       /* context name */
    Oid         ectxnamespace;  /* schema namespace */
    Oid         ectxowner;      /* owner role */
    Oid         ectxtypid;      /* referenced encoding type */

    Oid         ectxtargettypid;/* target SQL data type (0 / InvalidOid for any) */

#ifdef CATALOG_VARLEN
    text        ectxlabel;      /* human-readable version/schema label */
    bytea       ectxenctab;     /* raw encoding table / dictionary / schema / public key */
    bytea       ectxdectab;     /* raw decoding table / dictionary / schema / private key */
    text        ectxoptions[1]; /* options array (level, workers, strategy) */
#endif
} FormData_pg_encoding_context;

3.2 Catalog Conventions: enckind, enctype, and enclabel

PostgreSQL catalog conventions guide the naming and semantics of these fields:

Field Type Convention in Core Purpose
enctypkind char Matches relkind, prokind, aggkind High-level operational class: 'c' (compression), 's' (serialization), 'e' (encryption), 'q' (quantization), 't' (tokenization), 'd' (interning), 'm' (custom).
enctyptype NameData Matches srvtype, amtype The specific algorithm or engine format name (e.g. "zstd", "lz4", "avro", "huffman", "aes-gcm").
ectxtargettypid Oid Matches atttypid, opcintype, reltype Optional target SQL data type in pg_type (e.g. TEXTOID, JSONBOID, VECTOROID, or 0 for generic).
ectxlabel text Matches enumlabel, seclabel.label Human-readable tag, version string, or schema identifier (e.g. "v1.2", "avro:UserEvent.avsc").

3.3 Catalog Identifiers & Indexes

  • Relation ID: 9431 (EncodingContextRelationId)
  • Indexes:
    • EncodingContextNameNspIndexId (9432): Unique B-Tree index on (ectxname, ectxnamespace).
    • EncodingContextObjectIdIndexId (9433): Unique primary key B-Tree index on oid.
  • TOAST Table: Dedicated TOAST relation (9434 / 9435) for large enctab, dectab, and options.

4. In-Memory Runtime & Backend Cache

4.1 C Data Structures (src/include/access/encoding_context.h)

typedef struct EncodingContextData
{
    Oid         context_id;         /* Catalog OID in pg_encoding_context */
    Oid         type_id;            /* Catalog OID in pg_encoding_type */
    Oid         enc_fn;             /* Encode function OID (cached from pg_encoding_type) */
    Oid         dec_fn;             /* Decode function OID (cached from pg_encoding_type) */
    FmgrInfo    enc_flinfo;         /* Cached function call metadata */
    FmgrInfo    dec_flinfo;         /* Cached function call metadata */
    MemoryContext mcxt;             /* Owning memory context */

    /* Encoding Table / Dictionary / Schema */
    bytea      *enctab;             /* Raw encoding payload */
    int32       enctab_size;
    void       *enc_opaque;         /* Prepared runtime state (e.g. ZSTD_CDict*) */

    /* Decoding Table / Dictionary / Schema */
    bytea      *dectab;             /* Raw decoding payload (NULL => use enctab) */
    int32       dectab_size;
    void       *dec_opaque;         /* Prepared runtime state (e.g. ZSTD_DDict*) */
} EncodingContextData;

typedef EncodingContextData *EncodingContext;

4.2 Backend-Local Hash Table Cache

To eliminate per-tuple catalog overhead and repeated dictionary parsing:

  • Lifetime: Stored in a dedicated EncodingContextMemoryContext child of TopMemoryContext.
  • Key: Oid context_id.
  • SysCache Invalidation: Hooked into CacheRegisterSyscacheCallback for pg_encoding_context. Any ALTER or DROP triggers eviction and safe cleanup of prepared opaque objects.
  • Access Function: EncodingContext GetEncodingContext(Oid context_id).

5. Direct TOAST Integration (Patch 2)

5.1 Storage Layout

In create_toast_table(), Direct TOAST relations are defined with 6 columns:

  1. chunk_id (OID / OID8)
  2. chunk_seq (INT4)
  3. chunk_data (BYTEA)
  4. chunk_tids (TID[])
  5. chunk_tid_offsets (INT8[])
  6. chunk_ectx_id (OIDOID): The context_id associated with this chunk.

5.2 Slice Boundary Checkpoints

Root/Intermediate Chunk (chunk_tids, chunk_tid_offsets)
       │
       ├── Leaf Chunk 0 [Offset 0    .. 2047] (chunk_ectx_id = 16384) -> Decodes independently
       ├── Leaf Chunk 1 [Offset 2048 .. 4095] (chunk_ectx_id = 16384) -> Decodes independently
       └── Leaf Chunk 2 [Offset 4096 .. 6143] (chunk_ectx_id = 16384) -> Decodes independently

When fetching a slice:

  1. Binary search chunk_tid_offsets to locate the leaf chunk(s) covering [slice_offset, slice_offset + slice_len).
  2. Fetch the target chunk tuple directly by TID.
  3. Extract chunk_ectx_id from the chunk tuple.
  4. Lookup EncodingContext in the backend cache and invoke decfunc(context, chunk_data).
  5. Slice extraction is completely transparent to the decoder.

6. General Use Cases & Behavioral Examples

6.1 Schema-Driven Serialization (Apache Avro / Protobuf)

  • Problem: Storing JSONB carries repetitive field names across every single row.
  • Encoding Context:
    • enckind: 's'
    • enctype: "avro"
    • enctab: Avro writer schema JSON.
    • dectab: Avro reader schema with field projections.
  • Result: Compact binary representation with zero field-name storage overhead.

6.2 Write-Only Ingestion via Asymmetric Encryption

  • Problem: Web application tiers need to insert sensitive records (SSNs, credit cards, medical data) but should not possess the ability to decrypt them if compromised.
  • Encoding Context:
    • enckind: 'e'
    • enctype: "rsa-oaep"
    • enctab: RSA 4096-bit Public Key.
    • dectab: Private Key or KMS Decryption Handle (revoked from public/app roles).
  • Result: The app can execute encfunc and insert rows, but only privileged security backends can decode them.

6.3 Vector Quantization (Product Quantization for AI Embeddings)

  • Problem: 1536-dimensional float32 vectors require 6,144 bytes per row.
  • Encoding Context:
    • enckind: 'q'
    • enctype: "product_quantization"
    • enctab: Trained centroid codebook ($M=96$ subspaces, $K=256$ centroids).
    • dectab: Centroid lookup table.
  • Result: Vector reduced from 6,144 bytes to 96 bytes (bytea), achieving 98.4% space savings while retaining distance ranking properties.