You signed in with another tab or window. Reload to refresh your session.You signed out in another tab or window. Reload to refresh your session.You switched accounts on another tab or window. Reload to refresh your session.Dismiss alert
Triaged 2026-09-30. This was an AI-drafted ticket. The scope below replaces the original draft,
which is kept at the bottom. The full plan, with server behavior verified against dp-service main,
is in plan/tickets/16/plan.md.
Summary
Wrap the bucket-oriented v2 query, queryBuckets() (unary, resumable) and queryBucketsStream(), on the
existing QueryClient, taking the existing QueryParams. This is the only way to read back the array,
image, struct, and serialized columns that #17 made ingestable: querySamples is scalar-only. It is also
the first path on which ColumnMetadata (provenance) comes back live.
What triage changed
Most of the planned conversion layer already exists.interface to modernized annotation API #6 and Add ingestion API client (full surface: ingestData + streaming) #17 built data_frame_conversions
(per-column readers for all 14 typed kinds plus the legacy DataColumn, integer-nanosecond SamplingClock
expansion, pandas with narrow dtypes) and the frame-slicing code behind split_data_frame(). A DataBucket is one PV's column over one time axis, so the new work is: view a bucket as a one-column DataFrame, assemble a PV's buckets, and trim.
sampleStatusSelector is not "future". It shipped in 1.16.0, for samples only, and the server
rejects it on bucket queries. The client refuses it at request-build time with a ValueError.
limit counts buckets, not rows. Pages are also cut by bytes and still carry a token. Results are
sorted by PV then time, but the client does not rely on that order.
Scope
query_buckets(), iter_query_buckets(), iter_query_buckets_stream(), and QueryBucketsApiResult.
A new bucket_conversions module.
Pure Python: bucket_to_data_frame(), bucket_values() / bucket_timestamps() / bucket_column(), buckets_by_pv(), and an exact [begin, end)trim_bucket().
[analysis]: buckets_to_dataframes() returning one pandas DataFrame per PV, with opt-in
trimming and per-bucket metadata in attrs, plus a unary whole-query convenience.
Two PRs, as in #17: A is the client, the conversions, and unit tests. B is the live tests and docs.
Out of scope: queryData / queryTable / stats RPCs, useSerializedColumns=True, decoding serialized
payloads, a wide aligned pandas view, and a streaming pandas convenience.
Original AI-drafted ticket
Summary
Add Python client wrappers for the bucket-oriented v2 query RPCs on DpQueryService, complementing the sample-oriented client delivered in #7:
queryBuckets() — unary, one page, resumable via nextPageToken
queryBucketsStream() — server-streaming, fire-and-consume (no continuation tokens)
Part of the v2 query API epic (#10). #7 intentionally scoped itself to the sample-oriented path (querySamples() / querySamplesStream()) per its own ticket text, and was designed so the bucket methods slot in as an additive follow-on rather than a refactor.
Why this is a clean follow-on (not extra work deferred)
The request side is identical. QueryBucketsRequest wraps the exact same QuerySpec + ExecutionOptions + ResultRepresentation as QuerySamplesRequest. So #7's request-building machinery is reused verbatim:
the kind-neutral QueryParams dataclass
the PvQuery (PV) / ConfigQuery (CFG) criterion helpers
Key differences from the samples path that drive the work:
samplingClock expansion — a bucket can encode timestamps implicitly as startTime + periodNanos * i for count samples; the conversion layer must expand these into an explicit index.
16-arm typed-column dataValues oneof — distinct from the samples path's per-value DataValue oneof. Includes typed scalar columns, enum columns (enumId), image columns (imageDescriptor), struct columns (schemaId + bytes), and N-dimensional array columns (ArrayDimensions).
Whole / untrimmed boundary buckets — per the servicer docstring, queryBuckets() returns overlapping boundary buckets intact, so the first/last bucket may contain samples outside the requested [beginTime, endTime). (Contrast: querySamples() trims.) The conversion layer / docs must make this explicit; callers wanting strict trimming use the samples path.
Per-bucket-per-PV assembly — results are a list of DataBuckets (one PV's slice of time each), so assembling a coherent frame means grouping/concatenating across buckets, unlike the samples path's single aligned ColumnTable.
Conversion layer: DataBucket → pandas/NumPy (samplingClock expansion; 16-arm typed-column extraction; documented whole-bucket semantics), reusing the Phase-2 [analysis] optional-extra approach and DataValue/Image handling from interface to v2 time-series data query API #7 where applicable.
Unit tests per column arm + both dataTimestamps forms + untrimmed-boundary behavior; integration round-trip vs local :50052.
Independent of the Sample Status API work; the sampleStatusSelector addition (future) would benefit both sample and bucket paths equally since they share QuerySpec.
Summary
Wrap the bucket-oriented v2 query,
queryBuckets()(unary, resumable) andqueryBucketsStream(), on theexisting
QueryClient, taking the existingQueryParams. This is the only way to read back the array,image, struct, and serialized columns that #17 made ingestable:
querySamplesis scalar-only. It is alsothe first path on which
ColumnMetadata(provenance) comes back live.What triage changed
data_frame_conversions(per-column readers for all 14 typed kinds plus the legacy
DataColumn, integer-nanosecondSamplingClockexpansion, pandas with narrow dtypes) and the frame-slicing code behind
split_data_frame(). ADataBucketis one PV's column over one time axis, so the new work is: view a bucket as a one-columnDataFrame, assemble a PV's buckets, and trim.sampleStatusSelectoris not "future". It shipped in 1.16.0, for samples only, and the serverrejects it on bucket queries. The client refuses it at request-build time with a
ValueError.useSerializedColumnsdoes nothing on buckets, contrary to the proto. Typed columns always comeback typed, and only columns ingested serialized come back serialized. Serialized buckets are therefore
user data, not a deferred representation. They are passed through raw in pure Python and refused with a
clear error in pandas. Upstream doc fix: queryBuckets: correct useSerializedColumns docs, and document byte-cut pages, token scope, and result order dp-grpc#167.
limitcounts buckets, not rows. Pages are also cut by bytes and still carry a token. Results aresorted by PV then time, but the client does not rely on that order.
Scope
query_buckets(),iter_query_buckets(),iter_query_buckets_stream(), andQueryBucketsApiResult.bucket_conversionsmodule.bucket_to_data_frame(),bucket_values()/bucket_timestamps()/bucket_column(),buckets_by_pv(), and an exact[begin, end)trim_bucket().[analysis]:buckets_to_dataframes()returning one pandas DataFrame per PV, with opt-intrimming and per-bucket metadata in
attrs, plus a unary whole-query convenience.updates for every "not yet wrapped (interface to v2 bucket-oriented query API (queryBuckets / queryBucketsStream) #16)" pointer.
Two PRs, as in #17: A is the client, the conversions, and unit tests. B is the live tests and docs.
Out of scope:
queryData/queryTable/ stats RPCs,useSerializedColumns=True, decoding serializedpayloads, a wide aligned pandas view, and a streaming pandas convenience.
Original AI-drafted ticket
Summary
Add Python client wrappers for the bucket-oriented v2 query RPCs on
DpQueryService, complementing the sample-oriented client delivered in #7:queryBuckets()— unary, one page, resumable vianextPageTokenqueryBucketsStream()— server-streaming, fire-and-consume (no continuation tokens)Part of the v2 query API epic (#10). #7 intentionally scoped itself to the sample-oriented path (
querySamples()/querySamplesStream()) per its own ticket text, and was designed so the bucket methods slot in as an additive follow-on rather than a refactor.Why this is a clean follow-on (not extra work deferred)
The request side is identical.
QueryBucketsRequestwraps the exact sameQuerySpec+ExecutionOptions+ResultRepresentationasQuerySamplesRequest. So #7's request-building machinery is reused verbatim:QueryParamsdataclassPvQuery(PV) /ConfigQuery(CFG) criterion helpers_build_query_spec(params)seam (added in interface to v2 time-series data query API #7 specifically so_build_query_buckets_request()can reuse it)to_timestamp()time handling and page-token pagingThe bucket methods also live on the same
QueryClient/ query channel — no new client class.The real work: bucket response conversion
The bucket response is substantially richer than the samples
ColumnTable, and that's where this issue's effort concentrates:Key differences from the samples path that drive the work:
samplingClockexpansion — a bucket can encode timestamps implicitly asstartTime + periodNanos * iforcountsamples; the conversion layer must expand these into an explicit index.dataValuesoneof — distinct from the samples path's per-valueDataValueoneof. Includes typed scalar columns, enum columns (enumId), image columns (imageDescriptor), struct columns (schemaId+ bytes), and N-dimensional array columns (ArrayDimensions).queryBuckets()returns overlapping boundary buckets intact, so the first/last bucket may contain samples outside the requested[beginTime, endTime). (Contrast:querySamples()trims.) The conversion layer / docs must make this explicit; callers wanting strict trimming use the samples path.DataBuckets (one PV's slice of time each), so assembling a coherent frame means grouping/concatenating across buckets, unlike the samples path's single alignedColumnTable.Proposed scope
_build_query_buckets_request()(reusing_build_query_spec), unary + streaming_send_*,iter_query_buckets()/iter_query_buckets_stream(),QueryBucketsApiResult.DataBucket→ pandas/NumPy (samplingClock expansion; 16-arm typed-column extraction; documented whole-bucket semantics), reusing the Phase-2[analysis]optional-extra approach andDataValue/Imagehandling from interface to v2 time-series data query API #7 where applicable.dataTimestampsforms + untrimmed-boundary behavior; integration round-trip vs local :50052.serializedDataColumn) deferred, same as interface to v2 time-series data query API #7 (fail-loud).Dependencies
QueryParams/_build_query_spec/QueryClient).sampleStatusSelectoraddition (future) would benefit both sample and bucket paths equally since they shareQuerySpec.