Skip to content

[Feature] Support richer data types, starting with VECTOR<T, N> based on PIP-40 #197

Description

@ChaomingZhangCN

Search before asking

  • I searched in the issues and found nothing similar.

Motivation

Apache Paimon Java has introduced the VECTOR<T, N> data type based on PIP-40. It supports schema representation, regular data-file storage such as Parquet, dedicated vector storage, reads, writes, and Data Evolution.

Paimon C++ currently recognizes .vector. file names for some file-level bookkeeping, but it does not yet provide a VECTOR logical type, Arrow mapping, serialization, storage, or end-to-end read and write support.

This issue implements the VECTOR roadmap item tracked in #186.

Solution

Introduce VECTOR<T, N> support incrementally, while keeping the schema and storage behavior compatible with Apache Paimon Java.

Phase 1: Schema and regular Parquet storage

  • Add VECTOR<T, N> to the Paimon C++ logical type system.
  • Implement schema JSON serialization and deserialization compatible with Paimon Java.
  • Map VECTOR<T, N> to Arrow FixedSizeList<T, N>.
  • Support the element types defined by PIP-40:
    • BOOLEAN
    • TINYINT
    • SMALLINT
    • INT
    • BIGINT
    • FLOAT
    • DOUBLE
  • Validate that the dimension is positive and fixed.
  • Validate that the written vector length equals N.
  • Reject null vector elements.
  • Support reading and writing VECTOR columns in regular Parquet data files.
  • Add end-to-end append-table tests.
  • Add Java/C++ schema and file compatibility tests.

Phase 2: Data Evolution

  • Support adding and dropping VECTOR columns.
  • Support reading files written with previous table schemas.
  • Reject incompatible dimension changes, such as VECTOR<FLOAT, 3> to VECTOR<FLOAT, 5>.
  • Reject VECTOR columns as primary keys, partition keys, or sorting keys.
  • Add Data Evolution integration tests.

Phase 3: Dedicated vector storage

  • Support the vector file format configuration used by Paimon Java.
  • Support reading and writing dedicated *.vector.vortex files.
  • Integrate vector files with row tracking and Data Evolution.
  • Integrate vector files with scan planning, file commits, and conflict handling.
  • Add Java, Python, and C++ Vortex compatibility tests.

Initial scope

The initial implementation can focus on Phase 1, providing a usable end-to-end vertical slice through schema representation, Arrow mapping, and regular Parquet reads and writes.

Dedicated Vortex vector storage and Data Evolution can be delivered through follow-up pull requests under this issue.

The following items are not required for the initial implementation:

  • ORC VECTOR support
  • Vector indexes or similarity search
  • Changing vector dimensions through schema evolution
  • VECTOR values inside shared-shredding MAP columns
  • Element types not supported by Apache Paimon Java

PIP-40 should only be considered fully supported after all phases are complete. Completing Phase 1 means that regular Parquet VECTOR storage is supported, but does not imply support for dedicated Vortex vector files.

Anything else?

Related roadmap: #186

Are you willing to submit a PR?

  • I'm willing to submit a PR!

Metadata

Metadata

Labels

enhancementNew feature or request

Type

No type

Projects

No projects

Milestone

No milestone

Relationships

None yet

Development

No branches or pull requests

Issue actions