Class DictionaryChunkWriter

java.lang.Object
io.deephaven.extensions.barrage.chunk.BaseChunkWriter<Chunk<Values>>
io.deephaven.extensions.barrage.chunk.DictionaryChunkWriter
All Implemented Interfaces:
ChunkWriter<Chunk<Values>>

public class DictionaryChunkWriter extends BaseChunkWriter<Chunk<Values>>
Serializes a flat column chunk as Arrow Dictionary Encoded on the wire.

The RecordBatch for this column contains only the index buffer (an Int16/Int32/Int64 array mapping each row to a position in the dictionary). The dictionary values are shipped separately as DictionaryBatch messages by BarrageMessageWriterImpl using the DictionaryWriterState tracked by this writer.

Multiple columns may share the same dictionary id; they will all reference the same DictionaryWriterState instance (managed by a DictionaryWriterRegistry held on the stream view). A single DictionaryBatch is emitted per id per update, covering all new values introduced by any sharing column.

  • Constructor Details

  • Method Details

    • getDictId

      public long getDictId()
    • getValuesWriter

      public ChunkWriter<Chunk<Values>> getValuesWriter()
    • getValuesChunkType

      public ChunkType getValuesChunkType()
    • computeNullCount

      protected int computeNullCount(@NotNull ChunkWriter.Context context, @NotNull @NotNull RowSequence subset)
      Description copied from class: BaseChunkWriter
      Compute the number of nulls in the subset.
      Specified by:
      computeNullCount in class BaseChunkWriter<Chunk<Values>>
      Parameters:
      context - the context for the chunk
      subset - the subset of rows to consider
      Returns:
      the number of nulls in the subset
    • writeValidityBufferInternal

      protected void writeValidityBufferInternal(@NotNull ChunkWriter.Context context, @NotNull @NotNull RowSequence subset, @NotNull @NotNull BaseChunkWriter.SerContext serContext)
      Description copied from class: BaseChunkWriter
      Update the validity buffer for the subset.
      Specified by:
      writeValidityBufferInternal in class BaseChunkWriter<Chunk<Values>>
      Parameters:
      context - the context for the chunk
      subset - the subset of rows to consider
      serContext - the serialization context
    • getInputStream

      public ChunkWriter.DrainableColumn getInputStream(@NotNull ChunkWriter.Context context, @Nullable @Nullable RowSet subset, @NotNull @NotNull BarrageOptions options) throws IOException
      Get an input stream optionally position-space filtered using the provided RowSet.

      Always throws; callers must use getInputStream(Context, RowSet, BarrageOptions, DictionaryWriterState).

      Parameters:
      context - the chunk writer context holding the data to be drained to the client
      subset - if provided, is a position-space filter of source data
      options - options for writing to the stream
      Returns:
      a single-use DrainableColumn ready to be drained via grpc
      Throws:
      IOException
    • getInputStream

      public ChunkWriter.DrainableColumn getInputStream(@NotNull ChunkWriter.Context context, @Nullable @Nullable RowSet subset, @NotNull @NotNull BarrageOptions options, @NotNull @NotNull DictionaryWriterState state) throws IOException
      Builds the index-column stream for the current batch.

      Each logical row in subset is mapped to its 0-based dictionary index using state. Null rows produce a null-sentinel index value (so the returned stream's validity bitmap has 0 for those positions). Newly encountered non-null values are appended to the dictionary; callers can inspect DictionaryWriterState.hasDelta() after this call to determine whether a DictionaryBatch message must be emitted.

      Parameters:
      context - the chunk context holding the source data
      subset - the row offsets (positions within the source chunk) to include; null means all rows
      options - barrage serialization options
      state - the per-stream/per-id dictionary state; updated in place with any new values encountered
      Throws:
      IOException
    • getInputStream

      public ChunkWriter.DrainableColumn getInputStream(@NotNull ChunkWriter.Context context, @Nullable @Nullable RowSet subset, @NotNull @NotNull BarrageOptions options, @Nullable @Nullable DictionaryWriterRegistry dictionaryRegistry) throws IOException
      Get an input stream, threading a DictionaryWriterRegistry down to any dictionary-encoded writers nested within this writer (e.g. the values child of a run-end-encoded column). Writers that contain no dictionary encoding ignore the registry and behave identically to ChunkWriter.getInputStream(Context, RowSet, BarrageOptions).

      A DictionaryChunkWriter uses the registry to obtain (and register) its per-id DictionaryWriterState, so that the enclosing message writer emits the corresponding DictionaryBatch. Composite writers forward the registry to their children.

      Resolves this writer's DictionaryWriterState from dictionaryRegistry (registering it on first use) and builds the index stream. This is the entry point used when a dictionary-encoded column is nested inside another writer (e.g. the values child of a run-end-encoded column), where the enclosing writer forwards the registry rather than pre-resolving the state.

      The state is registered even for an empty batch (as long as a registry is supplied), so that the enclosing message writer emits an initial isDelta=false DictionaryBatch defining this field's id before the referencing RecordBatch, as strict Arrow consumers require. An empty batch adds no dictionary values, so that initial batch is empty. A null registry is only tolerated for an empty batch, which then produces the empty index stream directly.

      Parameters:
      context - the chunk writer context holding the data to be drained to the client
      subset - if provided, is a position-space filter of source data
      options - options for writing to the stream
      dictionaryRegistry - the per-stream dictionary registry; may be null when the caller does not support dictionary encoding (a nested dictionary-encoded writer will then throw)
      Returns:
      a single-use DrainableColumn ready to be drained via grpc
      Throws:
      IOException
    • getEmptyIndexStream

      public ChunkWriter.DrainableColumn getEmptyIndexStream(@NotNull @NotNull BarrageOptions options) throws IOException
      Returns the ChunkWriter.DrainableColumn for an empty (0-row) dictionary-encoded column batch. The column's validity and index buffers are both empty. This does not touch the DictionaryWriterState; callers register the state separately (see getInputStream(Context, RowSet, BarrageOptions, DictionaryWriterRegistry)) when an initial empty DictionaryBatch is required for this id.
      Throws:
      IOException