diff --git a/CLAUDE.md b/CLAUDE.md
index 7f3292d..a9bf749 100644
--- a/CLAUDE.md
+++ b/CLAUDE.md
@@ -231,7 +231,7 @@ plan documents one change, `CLAUDE.md` documents the invariant it established.
## Key Files
- `src/dp_python_lib/client/mldp_client.py` - Main client wrapper for the gRPC services
-- `src/dp_python_lib/client/ingestion_client.py` - Ingestion service client with methods like `register_provider()`
+- `src/dp_python_lib/client/ingestion_client.py` - Ingestion service client (issue #17; `plan/tickets/17/plan.md`): `register_provider()` (its result now exposes `provider_id` / `is_new_provider`), `ingest_data()` (unary), `ingest_data_stream()` (client-streaming, one summary), `iter_ingest_data_bidi_stream()` (one ack or reject per request), `query_request_status()` with the `RequestStatusQuery` (`RS`) helpers and the `IngestionRequestStatus` enum, and `await_request_statuses()`. `IngestDataRequestParams` takes a `common.DataFrame` from the `data_frame` builders (no separate ingestion column model), re-validates it with `validate_data_frame()`, defaults `client_request_id` to a uuid4, and caps both ids at `MAX_BUDGETED_ID_CHARS`; `chunked_request_params()` names `split_data_frame()` chunks `-`. See the Ingestion API section below for the server behaviors it encodes
- `src/dp_python_lib/client/annotation_client.py` - Annotation service facade; groups feature-scoped clients sharing the one `DpAnnotationService` channel (`.pv_metadata`, `.machine_config`, `.sample_status`, `.datasets`, `.annotations`, `.export` — every implemented `DpAnnotationService` feature area)
- `src/dp_python_lib/client/pv_metadata_client.py` - PV metadata client (`save_pv_metadata()`, `get_pv_metadata()`, `query_pv_metadata()`, `iter_pv_metadata()`, `delete_pv_metadata()`) plus the `PvMetadataQuery` (`Q`) criterion helpers
- `src/dp_python_lib/client/machine_config_client.py` - Machine configuration client covering both configurations (`save_configuration()`, `get_configuration()`, `query_configurations()`, `iter_configurations()`, `delete_configuration()`) and their temporal activations (`save_configuration_activation()`, `get_configuration_activation()`, `query_configuration_activations()`, `iter_configuration_activations()`, `delete_configuration_activation()`, `get_active_configurations()`). Includes the `ConfigurationQuery` (`C`) and `ConfigurationActivationQuery` (`CA`) criterion helpers. The shared time converters it used to own now live in `time_conversions.py`. Get/delete activation take a composite key (`client_activation_id` XOR `configuration_name`+`start_time`). Activation `end_time` is optional — omit it for an open-ended activation ("still in effect"); the field is then genuinely absent on the wire. Read it back with the module-level `activation_is_open()` / `activation_end_time()` (#26), which work on an activation from any read path: reading `.endTime` directly on an open record silently yields a zero `Timestamp` (1970), and passing that back as `end_time=` in a re-save gets the save rejected. `activation_end_time()` returns `None` for an open record, so `end_time=activation_end_time(current)` is the correct carry-forward. Presence is the only test: an `endTime` present with value 0 is closed, not open
@@ -240,15 +240,16 @@ plan documents one change, `CLAUDE.md` documents the invariant it established.
- `src/dp_python_lib/client/dataset_client.py` - DataSet client (`save_dataset()`, `get_dataset()`, `query_datasets()`, `iter_datasets()`, `delete_dataset()`, plus the `get_datasets(ids)` batch fetch that avoids the annotation-listing N+1) with the `DataSetQuery` (`DS`) criterion helpers and the `data_block()` builder. `data_block()` is the only place `begin < end` is checked — the server does not
- `src/dp_python_lib/client/annotations_client.py` - Annotations client (`save_annotation()`, `get_annotation()`, `query_annotations()`, `iter_annotations()`, `delete_annotation()`, `get_calculations()`) with the `AnnotationQuery` (`AQ`) criterion helpers and the `calculations()` builder, which takes a `dict[str, DataFrame]` so frame-name uniqueness is true by construction. Note `AnnotationsClient` (feature client) vs `AnnotationClient` (facade)
- `src/dp_python_lib/client/export_client.py` - Export client (`export_data()`) with the `ExportFormat` str enum and the `calculations_spec()` builder
-- `src/dp_python_lib/client/time_conversions.py` - The two shared time converters and the `TimestampInput` alias: `to_timestamp()` (tz-aware datetime / epoch seconds / `common.Timestamp` → `Timestamp`; naive datetimes raise) and its inverse `to_epoch_nanos()` (`Timestamp` → integer epoch nanoseconds), plus `NANOS_PER_SECOND`. **The datetime path uses integer arithmetic, never `datetime.timestamp()`** — that returns a float64, which cannot hold present-day epoch seconds at sub-microsecond resolution and moved 99.7% of microsecond datetimes by up to ~119ns, breaking the exact-match contract sample status and provenance depend on. Both were defined in `machine_config_client` as the first module to need them, and six others grew imports from there — which read as though datasets, queries, and DataFrames depended on the machine configuration API; `to_epoch_nanos()` had also been written out privately three separate times. **A leaf module**: it imports only stdlib and the generated protos, so any client module can use it without an import cycle. New time conversions belong here
-- `src/dp_python_lib/client/data_frame.py` - Builders for `common.DataFrame`, the shared time-series payload (also ingestion's `ingestionDataFrame`, so #17 extends this rather than forking it): the `sampling_clock()` / `timestamp_list()` / `timestamp_count()` axis helpers **relocated here from `sample_status_client`** (and re-exported from it, so existing imports keep working), the typed scalar column builders (`double_column`, `float_column`, `int64_column`, `int32_column`, `bool_column`, `string_column`, `enum_column`), the legacy `data_column()` escape hatch (a `None` entry becomes an unset oneof — the only way to express a gap on a shared axis), the provenance helpers (`column_metadata`, `provenance`, `pv_source`, `calculations_source`), and `data_frame()` assembly, which routes columns by type and enforces the server's **shape** rules client-side (non-blank names, non-empty values, count match, name uniqueness across all types) while leaving its size caps server-side. Column names must be **non-blank**, not merely non-empty, and `timestamp_count()` validates a hand-built axis the way the builders do: a `SamplingClock` needs a positive `periodNanos` as well as a non-zero count, and a `TimestampList` must be **strictly increasing** (duplicates included — two samples cannot claim one instant). The read path rejects all of these, and anything `data_frame()` accepts must be readable back. Shared with `SampleStatusFrame`, which validates the same way. Sample counts come from the right field per kind: `dataValues` for a `DataColumn`, `images` for an `ImageColumn` (which has no `values` field at all), and `len(values) / prod(dims)` for an array column. An array column's sample count is `len(values) / prod(dims)`, and a value count that is not a **whole multiple** of `prod(dims)` is rejected rather than floor-divided into a passing count — the read path applies the same rule, so a frame this accepts is always one `data_frame_conversions` can read back. `data_column()` maps integers and floats by `numbers.Integral` / `numbers.Real` (and NumPy's bool by type name, without importing NumPy), so NumPy scalars map like their Python counterparts: `np.float64` subclasses `float` but `np.int64` and `np.bool_` subclass nothing here, and matching on exact Python type would accept some and reject others. Array/image/struct/serialized builders are #17's; hand-built ones pass through
+- `src/dp_python_lib/client/time_conversions.py` - The two shared time converters and the `TimestampInput` alias: `to_timestamp()` (tz-aware datetime / epoch seconds / `common.Timestamp` → `Timestamp`; naive datetimes raise) and its inverse `to_epoch_nanos()` (`Timestamp` → integer epoch nanoseconds), plus `NANOS_PER_SECOND` and `from_epoch_nanos()` (integer epoch nanoseconds → `Timestamp`; it replaced private copies in `data_frame_conversions` and `sample_status_conversions` once `split_data_frame()` became a third caller). **The datetime path uses integer arithmetic, never `datetime.timestamp()`** — that returns a float64, which cannot hold present-day epoch seconds at sub-microsecond resolution and moved 99.7% of microsecond datetimes by up to ~119ns, breaking the exact-match contract sample status and provenance depend on. Both were defined in `machine_config_client` as the first module to need them, and six others grew imports from there — which read as though datasets, queries, and DataFrames depended on the machine configuration API; `to_epoch_nanos()` had also been written out privately three separate times. **A leaf module**: it imports only stdlib and the generated protos, so any client module can use it without an import cycle. New time conversions belong here
+- `src/dp_python_lib/client/data_frame.py` - Builders for `common.DataFrame`, the shared time-series payload (also ingestion's `ingestionDataFrame`, so #17 extends this rather than forking it): the `sampling_clock()` / `timestamp_list()` / `timestamp_count()` axis helpers **relocated here from `sample_status_client`** (and re-exported from it, so existing imports keep working), the typed scalar column builders (`double_column`, `float_column`, `int64_column`, `int32_column`, `bool_column`, `string_column`, `enum_column`), the legacy `data_column()` escape hatch (a `None` entry becomes an unset oneof — the only way to express a gap on a shared axis), the provenance helpers (`column_metadata`, `provenance`, `pv_source`, `calculations_source`), and `data_frame()` assembly, which routes columns by type and enforces the server's **shape** rules client-side (non-blank names, non-empty values, count match, name uniqueness across all types) while leaving its size caps server-side. Column names must be **non-blank**, not merely non-empty, and `timestamp_count()` validates a hand-built axis the way the builders do: a `SamplingClock` needs a positive `periodNanos` as well as a non-zero count, and a `TimestampList` must be **strictly increasing** (duplicates included — two samples cannot claim one instant). The read path rejects all of these, and anything `data_frame()` accepts must be readable back. Shared with `SampleStatusFrame`, which validates the same way. Sample counts come from the right field per kind: `dataValues` for a `DataColumn`, `images` for an `ImageColumn` (which has no `values` field at all), and `len(values) / prod(dims)` for an array column. An array column's sample count is `len(values) / prod(dims)`, and a value count that is not a **whole multiple** of `prod(dims)` is rejected rather than floor-divided into a passing count — the read path applies the same rule, so a frame this accepts is always one `data_frame_conversions` can read back. `data_column()` maps integers and floats by `numbers.Integral` / `numbers.Real` (and NumPy's bool by type name, without importing NumPy), so NumPy scalars map like their Python counterparts: `np.float64` subclasses `float` but `np.int64` and `np.bool_` subclass nothing here, and matching on exact Python type would accept some and reject others. #17 added the non-scalar builders: `double_array_column` / `float_array_column` / `int32_array_column` / `int64_array_column` / `bool_array_column` (samples as nested sequences or NumPy arrays, duck-typed; shape inferred and required identical across samples; flattened **row-major**, which is what the read path assumes; 1–3 dims; explicit `dims=` may shape flat samples or restate a shape, never reinterpret one), `image_column` (one descriptor per column, as in the proto), `struct_column`, and `serialized_column`. **Each kind's structural fields are checked in `_check_column()`, not only in the builders** — an enum's `enumId`, 1–3 array dims, an image's descriptor (positive width/height/channels, non-blank encoding), a struct's `schemaId`, a serialized column's `encoding` — because a hand-built column bypasses every builder and the server rejects all of these (the recurring #6 defect shape). `validate_data_frame(frame)` applies the same checks to an assembled frame (ingestion re-validates with it; its `caller=` keyword names the entry point an error message starts with, so a frame rejected by `IngestDataRequestParams` or `split_data_frame()` is not reported as a `data_frame()` error); `data_frame()` shares the column-list check rather than calling it, so messages keep the caller's index. `split_data_frame(frame, *, max_rows, max_bytes, max_span_nanos)` chunks lazily along the time axis: `max_bytes` bounds the **whole `IngestDataRequest`**, with room for two worst-case ids (`MAX_BUDGETED_ID_CHARS` = 256 chars × 4 UTF-8 bytes, computed from the message, not hard-coded); `max_span_nanos` is first-to-last timestamp, the server's bucket-span measure; `SamplingClock` chunk starts are integer nanoseconds. No limit defaults: `SERVER_DEFAULT_MAX_MESSAGE_BYTES` (4,096,000) is exported as a reference only, since the server value is deployment configuration. A frame with a `SerializedDataColumn` cannot be split
- `src/dp_python_lib/client/data_frame_conversions.py` - Reading a `DataFrame` back. Pure Python (no extras): `data_frame_timestamps()` (integer-nanosecond axis expansion), `column_values()` (standalone per-column converter, written so the bucket query #16 can reuse it; it yields exactly one entry per sample for every column kind, with the structural fields a payload cannot be interpreted without kept in companion accessors: array dims via `column_dimensions()` / `data_frame_column_dimensions()` so `[2,2]` and `[4]` stay distinguishable, an image's width/height/channels/encoding via `image_descriptor_dict()` / `data_frame_image_descriptors()`, and a struct's `schemaId` via `column_schema_id()` / `data_frame_schema_ids()`), `data_frame_columns()`, `column_metadata_dict()`. An axis that is set but empty is rejected rather than converted to a zero-row table. Behind `[analysis]`: `data_frame_to_pandas()` (UTC index built from int64 nanos, `ColumnMetadata` in `df.attrs`; each Series carries the **narrow dtype its column type implies** — `float32`/`int32`, not pandas' widened inference), `data_frame_from_pandas()` (dtype→typed column; a NaN anywhere is fail-loud, since a dense typed column cannot express a gap; rebuilds each column's `ColumnMetadata` from `df.attrs` via `column_metadata_from_dict()`). **A frame round-trips through pandas to byte equality** apart from the deliberate `SamplingClock`→`TimestampList` axis change: column types and provenance both survive. An `EnumColumn`'s codes are int32 and indistinguishable from a plain `Int32Column` by dtype, so its `enumId` rides in `df.attrs["enum_ids"]` — carried even under `exclude_column_metadata=True`, since without it the column cannot be rebuilt as an enum at all, and the `calculations_to_dataframes()` / `calculations_from_dataframes()` bridges. The pandas direction always emits a `TimestampList`, never an inferred `SamplingClock`; reads each instant as `Timestamp.value` (**always nanoseconds**, unlike a raw int64 view, which is in the index's own storage unit — a `datetime64[us]` index viewed as int64 lands 1000× too early); rejects `NaT`, whose integer form is a valid-looking instant; and rejects duplicate column names up front (a duplicated label makes `df[name]` a DataFrame, which would otherwise fail deep inside as pandas' "truth value of a Series is ambiguous"). `column_metadata_dict()` reports **only the origin arm actually set** on each provenance source — a PV source has no `calculations_column` key and vice versa — plus `time_range` as epoch nanoseconds when present; an absent range has no key rather than a fabricated `(0, 0)`. A pandas round trip preserves every value, dtype, and timestamp but **not column order**: a `DataFrame` stores each column kind in its own repeated field, so columns come back grouped by type
- `src/dp_python_lib/client/query_client.py` - v2 time-series query client (sample-oriented) exposed as `client.query`. Low-level wrappers `query_samples()` (unary, one resumable page) and `iter_query_samples()` (transparent paging), plus `iter_query_samples_stream()` (server-streaming, fire-and-consume, lazy). Queries are described by a kind-neutral `QueryParams` built from the `PvQuery` (`PV`) and `ConfigQuery` (`CFG`) criterion helpers; shares a `_build_query_spec()` seam so a future bucket request builder reuses it. Results wrap the raw `ColumnTable` (`.column_table`, `.next_page_token`); `.to_dataframe()`/`.to_numpy()` delegate to `query_conversions` (Phase 2, optional `[analysis]` extra)
- `src/dp_python_lib/client/query_conversions.py` - Pythonic conversions for query results (optional `[analysis]` extra: pandas/numpy/openpyxl, imported lazily). `data_value_to_python()` (oneof extractor: scalars→native, timestamp→epoch-nanos, array→list, structure→dict, image→`Image` wrapper, fail-loud on unhandled arm), `column_table_to_dataframe()` (UTC datetime index + one column per DataColumn; dense-alignment and duplicate-column-name fail-loud; ColumnMetadata in `df.attrs`), `column_table_to_numpy()` (dict of 1-D arrays; complex arms stay 1-D object arrays rather than collapsing to 2-D), `dataframe_to_excel()` (thin `to_excel()` wrapper: row-limit guard, tz-drop, complex-cell stringification), and `query_samples_to_dataframe()`/`stream_query_samples_to_dataframes()` whole-query conveniences (unary concats by column name; streaming yields per-page frames lazily)
- `src/dp_python_lib/client/service_api_client_base.py` - Base class for the service clients: owns the channel and the one-per-client gRPC stub, and provides `_dispatch()`, the shared three-tier sender that all 18 unary `_send_*` methods delegate to
- `src/dp_python_lib/client/query_support.py` - Helpers shared by the criteria-based query clients, currently `check_at_most_one_text_criterion()` (Mongo cannot AND two `$text` clauses). It is generic over the different criterion types because each names its oneof `criterion` and its full-text arm `textCriterion`. New shared query helpers belong here rather than in whichever feature client happened to need one first — importing a private name across feature modules makes the importing module's dependencies misleading
- `tests/unit/test_service_api_client_base.py` - Unit tests for `_dispatch` itself (success, business error, unrecognized response, `RpcError` with and without a resolvable `code()`, unexpected exception, and the `request_log`/`success_log` hooks)
-- `tests/unit/test_ingestion_client.py` - Unit tests for IngestionClient functionality
+- `tests/unit/test_ingestion_client.py` - Unit tests for IngestionClient: registration, request params (id generation, length cap, frame re-validation), `chunked_request_params()`, unary/stream/bidi result handling with mocked stubs, the size-gated `RESOURCE_EXHAUSTED` hints, `RS` helpers, and `await_request_statuses()` against a fake clock (one query per poll, the skew-adjusted floor, timeout naming the missing ids)
+- `tests/unit/test_ingestion_streaming_grpc.py` - The streaming ingest paths against an **in-process grpcio server** on an ephemeral localhost port, because a mocked stub cannot reproduce grpcio consuming the request iterator on its own thread: the caller's own exception re-raised in place of "Exception iterating requests!", and closing a bidi stream's iterator (directly or via `contextlib.closing` around a `break`) cancelling the call. Also pins the grpcio premise itself, so a grpcio that starts chaining the exception fails loudly
- `tests/unit/test_pv_metadata_client.py` - Unit tests for PvMetadataClient functionality
- `tests/unit/test_machine_config_client.py` - Unit tests for the Configuration side of MachineConfigClient
- `tests/unit/test_machine_config_activation_client.py` - Unit tests for the ConfigurationActivation side of MachineConfigClient (incl. composite-key validation, timestamp handling, getActiveConfigurations)
@@ -260,7 +261,7 @@ plan documents one change, `CLAUDE.md` documents the invariant it established.
- `tests/unit/test_annotation_client.py` - Unit tests pinning the `AnnotationClient` facade wiring (every feature client present, one shared channel, one stub apiece)
- `tests/integration/test_datasets_annotations_integration.py` - Live-server round trip for datasets/annotations/calculations; ingests its own samples first, because `saveDataSet` requires archived PVs
- `tests/integration/test_query_helper_relaxations_integration.py` - Live-server coverage for the #40 key-only `attributes()` search and the #41 browse-all `criteria`, on PV metadata, configurations, activations, and the v2 `PvQuery.attr` selector. Both rest on server behavior a unit test cannot reach: a unit test asserts the request carries `values == []`, but only a real server distinguishes an existence filter from an `$in: []` that matches nothing. Each test therefore stores an attribute value, asserts the key-only form finds the record, and asserts a query for a *different* value does not -- that pairing is what makes the first assertion meaningful. The v2 class ingests its own samples (the selector needs archived data) and polls for bucket visibility rather than sleeping; it establishes that visibility with a *name-list* selector before asserting the negative case, so an empty result can only mean the attribute selector matched nothing. Catalogue records are torn down per run, but the ingested samples are not -- the archive has no delete RPC, so each run leaves a few samples under a run-unique PV name, the same residue `test_datasets_annotations_integration.py` leaves
-- `tests/unit/test_time_conversions.py` - Unit tests for the shared time converters (`to_timestamp()` input forms and the naive-datetime/bool/unsupported-type rejections; `to_epoch_nanos()` exactness and its round trip with `to_timestamp()`)
+- `tests/unit/test_time_conversions.py` - Unit tests for the shared time converters (`to_timestamp()` input forms and the naive-datetime/bool/unsupported-type rejections; `to_epoch_nanos()` exactness and its round trip with `to_timestamp()`; `from_epoch_nanos()` and its round trip)
- `tests/unit/test_data_frame.py` - Unit tests for the data_frame builders (axis relocation, each typed column, `data_column()` bool-before-int and unset-oneof handling, provenance helpers, and every `data_frame()` shape rule incl. array dims and serialized-column name-only checks)
- `tests/unit/test_data_frame_conversions.py` - Unit tests for data_frame_conversions (nanosecond-exact expansion, per-column conversion incl. array reshaping, duplicate-name fail-loud, and — skipping cleanly without `[analysis]` — the pandas round trip, dtype mapping, and NaN fail-loud)
- `tests/unit/test_query_client.py` - Unit tests for QueryClient (request building, three-tier error handling, unary paging, streaming, `PvQuery`/`ConfigQuery` helpers, `QueryParams` validation)
@@ -303,10 +304,19 @@ plan documents one change, `CLAUDE.md` documents the invariant it established.
`"Calling API"` / `" completed successfully"`. Pass one whenever the message names the
entity or reports a count off the response — that detail is the reason the parameters exist, and
nothing in the test suite asserts on log content, so a dropped message fails silently.
+ `rpc_error_hint` (optional, #17) is passed the caught `RpcError` and returns text to **append** to
+ `"gRPC error: "` -- never substitute, since that prefix is contract. The ingestion senders use
+ it to point an oversized request at `split_data_frame()`.
- **The server-streaming senders deliberately do not use `_dispatch`.**
`_send_query_samples_stream()` and `_send_query_sample_statuses_stream()` *yield* one result per
streamed message — error results included, for the public `iter_*` wrapper to convert into a
`RuntimeError` — which is a different contract from returning a single result. Leave them as they are.
+- **Neither do the two client-side streaming ingest senders** (#17), each for its own reason.
+ `_send_ingest_data_stream()` returns a single result but must **keep the response on error**: a partial
+ reject arrives as an `ExceptionalResult` whose `rejectedRequestIds` is the only record of which requests
+ failed, and `_dispatch` drops the response on error. `_send_ingest_data_bidi_stream()` yields like the query
+ stream senders, except that a per-request reject is yielded *with* its response, and only a result without one
+ (a transport error) makes the public wrapper raise.
- Where one gRPC service backs several feature areas (e.g. `DpAnnotationService` covers PV metadata,
machine configuration, and annotations), use a lightweight facade (`AnnotationClient`) that owns the
shared channel and exposes feature-scoped clients as attributes (`annotation.pv_metadata`). This
@@ -457,6 +467,95 @@ channel = grpc.insecure_channel("localhost:50051")
client = MldpClient(ingestion_channel=channel)
```
+### Ingestion API (Ingestion Service)
+
+Issue #17 (`plan/tickets/17/plan.md`) wrapped the ingestion side at `client.ingestion_client` (the attribute
+keeps its name; a `client.ingestion` alias was declined as a cross-cutting rename). The payload is the same
+`common.DataFrame` as an annotation's calculations, built with the `data_frame` builders.
+
+```python
+from datetime import datetime, timezone
+from dp_python_lib.client import (
+ MldpClient, RegisterProviderRequestParams, IngestDataRequestParams, IngestionRequestStatus,
+ chunked_request_params,
+)
+from dp_python_lib.client import data_frame as dfb
+
+ingestion = MldpClient().ingestion_client
+registration = ingestion.register_provider(
+ RegisterProviderRequestParams("bpm-daq", description=None, tag_list=None, attribute_map=None))
+provider_id = registration.provider_id # None if registration failed
+assert provider_id is not None, registration.result_status.message
+
+t0 = datetime(2026, 9, 30, 18, tzinfo=timezone.utc)
+frame = dfb.data_frame(dfb.sampling_clock(t0, period_nanos=1_000_000, count=3),
+ [dfb.double_column("BPMS:GUNB:314:X", [0.1, 0.2, 0.3])])
+
+params = IngestDataRequestParams(provider_id, frame) # client_request_id defaults to a uuid4
+since = datetime.now(timezone.utc) # captured BEFORE sending, for the status query
+ack = ingestion.ingest_data(params) # is_error on a reject; an ack is NOT ingestion
+statuses = ingestion.await_request_statuses(provider_id, [params.client_request_id], since=since)
+ok = all(d.ingestionRequestStatus == IngestionRequestStatus.SUCCESS
+ for docs in statuses.values() for d in docs)
+
+# a frame too big for one message: chunk it and stream the chunks lazily
+chunks = dfb.split_data_frame(frame, max_bytes=dfb.SERVER_DEFAULT_MAX_MESSAGE_BYTES)
+summary = ingestion.ingest_data_stream(chunked_request_params(provider_id, chunks))
+print(summary.num_requests, summary.rejected_request_ids)
+```
+
+Invariants worth knowing before touching this code (dp-service citations are in the plan's T-findings):
+
+- **An ack means "passed validation", not "ingested".** The service validates, acks or rejects, and only then
+ queues the request. An **unknown `providerId` is acked** and then fails asynchronously -- contrary to two proto
+ comments (filed as osprey-dcs/dp-grpc#165). Bucket-building errors and Mongo insert failures are equally
+ invisible at ack time, including a re-ingest of the same PV with the same first timestamp, which fails as a
+ duplicate bucket `_id` and can leave **a partial write whose status says ERROR** (`insertMany` is ordered). The
+ request-status document is written *after* the buckets, so SUCCESS means persisted and queryable, and it is the
+ only confirmation there is.
+- **`IngestionRequestStatus`'s zero value is SUCCESS**, the same silent-zero trap as an activation's `endTime`
+ (#26): an unset status reads as success. Nothing in the client defaults a status, and a *missing* document is
+ "unknown", never success -- a few server paths write none at all (a reject while shutting down, a handler that
+ throws), which is why `await_request_statuses()` has a timeout.
+- **`queryRequestStatus` cannot ask about several request ids at once** (`RequestIdCriterion` holds one, and
+ criteria AND) **and has no paging** (every match in one message; osprey-dcs/dp-grpc#165 and
+ osprey-dcs/dp-service#302 add it). So `await_request_statuses()` issues **one** query per poll -- provider plus
+ a time range from `since` -- and matches ids client-side. `since` is required and has no default: the time floor
+ bounds the unpaged response, and it keeps out documents from earlier runs that reused an id, since the server
+ enforces no uniqueness. The floor is backed off by `REQUEST_STATUS_CLOCK_SKEW` (60 s) because `createdAt` is the
+ server's clock; the end is left unset so "now" is the server's too. The time range matches `createdAt` -- when
+ the job *finished* -- inclusive at both ends, not the receipt time the proto implies.
+- **`ingest_data_stream()` keeps the response on error** (see the Client Implementation Pattern): a partial reject
+ is `is_error` with `rejected_request_ids`, and it does **not** raise, because the other requests were accepted.
+ `num_requests` is then `None`, since the server omits it.
+- **`iter_ingest_data_bidi_stream()` yields a reject rather than raising** -- deliberately unlike
+ `iter_query_samples_stream()`: a reject is a fact about one request, and the server keeps going. A transport
+ error raises `RuntimeError`.
+- **grpcio hides an exception raised by a request iterator**, reporting only UNKNOWN "Exception iterating
+ requests!". `_RequestFeed` records it and the sender re-raises the caller's own exception, chained from the
+ `RpcError`, with a sent-count note (PEP 678, so Python 3.11+) and a log line. **The count is an upper bound,
+ not a delivery receipt**: grpcio's cancellation races the sends, and the in-process tests saw a request handed to
+ grpcio never reach the server. Say "any that reached the server", never "were ingested".
+- **Closing a bidi stream's iterator cancels the call; a bare `break` does not.** grpcio pulls requests on its own
+ thread; without an explicit `cancel()` it keeps *sending* after the caller stops reading (an in-process test saw
+ all 1,000 queued requests drain into the server). The sender cancels in its `finally`, and the public wrapper
+ closes the sender with `contextlib.closing`, so closing the returned iterator cancels at once. But a `break` out
+ of a loop over a generator that is still referenced does not close it -- that waits for garbage collection -- so
+ the documented way to stop early is the caller's own `with contextlib.closing(...)`, and the in-process tests pin
+ exactly that. Do not describe the cancel as happening on "abandoning" or "breaking out of" the stream.
+- **The client is stricter than the server** on whitespace-only column names and duplicate timestamps (both
+ accepted by ingestion), because the library must read back what it writes; a caller who needs either can call
+ the stub directly. It is *not* stricter on caps: size limits stay server-side, except that `split_data_frame()`
+ lets the caller name them.
+- **Over the inbound message limit (4,096,000 bytes by default) a call fails with `RESOURCE_EXHAUSTED`**, not an
+ `ExceptionalResult`, and on a stream it kills every request after it. The hint that points at
+ `split_data_frame()` is gated on the size-violation *details* text, because `RESOURCE_EXHAUSTED` also covers
+ quotas; both matched texts live in module constants, since neither is a stable API.
+- **A frame's column provenance survives ingestion only into a 1.16.0 or later server**; the rest of the
+ ingestion proto is unchanged since rel-1.15.0.
+- **Non-scalar columns cannot be read back live yet**: `querySamples` is scalar-only, so array, image, struct,
+ and serialized data is verified only as far as ingestion (ack + SUCCESS) until the bucket query (#16).
+
### PV Metadata API (Annotation Service)
PV metadata methods are exposed under the `annotation` facade at `client.annotation.pv_metadata`
(available whenever an annotation channel/config is provided):
@@ -686,8 +785,8 @@ Invariants worth knowing before touching this code:
validated client-side (the client cannot know what is archived), but any test or example must use an archived PV.
`ingestData()` acks *before* the bucket becomes queryable, so a `saveDataSet` issued immediately after ingesting
still fails; `tests/integration/test_datasets_annotations_integration.py` probes until it succeeds rather than
- sleeping a fixed interval, and ingests through the generated stub because `IngestionClient` wraps only
- `registerProvider()` until #17.
+ sleeping a fixed interval, and ingests through the generated stub because it predates the ingestion client;
+ #17's second PR moves it onto `ingest_data()` and `await_request_statuses()`.
- **The server does not check `begin < end` on a `DataBlock`** (it checks only that each bound is non-zero and that
`pvNames` is non-empty, and never compares the two bounds), so `data_block()`'s check is the only one there is.
- **A `DataBlock`'s range is half-open, `[begin, end)`** — the same convention as the v2 query API's `QueryParams`,
diff --git a/README.md b/README.md
index 65ee8b6..9214fed 100644
--- a/README.md
+++ b/README.md
@@ -88,8 +88,13 @@ for their service.
bridges to pandas under the optional `[analysis]` extra.
- **Export** — `client.annotation.export`. Export a saved dataset, ad-hoc data blocks, and/or
calculations to HDF5, CSV, or XLSX. The file is written on the server; there is no retrieval RPC.
-- **Provider registration** — `client.ingestion_client.register_provider()`. The rest of the
- ingestion API is not yet implemented.
+- **Ingestion** — `client.ingestion_client`. Register a provider with `register_provider()`, then
+ ingest a `data_frame` frame with `ingest_data()`, `ingest_data_stream()` (many requests, one
+ summary), or `iter_ingest_data_bidi_stream()` (an ack or reject per request). An ack means only
+ that a request passed validation: confirm it landed with `await_request_statuses()` or
+ `query_request_status()`. `data_frame.split_data_frame()` cuts a large frame into chunks under
+ the server's message-size and time-span limits, and the `data_frame` builders now cover array,
+ image, struct, and serialized columns as well as scalars.
**Supporting framework:** YAML + environment-variable configuration (`MLDP_*`, via
pydantic-settings), TLS-capable channel creation, hierarchical logging, three-tier error handling
@@ -105,10 +110,6 @@ older than your `dp_python_lib` will not implement everything listed here. The
**Low-level API coverage**
- **Ingestion Service**
- - `ingestData()` / `ingestDataStream()` / `ingestDataBidiStream()` — full ingestion client with a
- shared DataFrame payload model ([issue #17](https://github.com/osprey-dcs/dp-python-lib/issues/17);
- also unblocks the closed-loop query integration test and the cookbook's ingestion recipe)
- - `queryRequestStatus()` — async status of ingestion requests
- `subscribeData()` — receive data for specified PVs from the ingestion stream
- **Query Service**
- `queryBuckets()` / `queryBucketsStream()` — raw data buckets
diff --git a/doc/release-notes/NEXT.md b/doc/release-notes/NEXT.md
index b0ab122..6bd3cdc 100644
--- a/doc/release-notes/NEXT.md
+++ b/doc/release-notes/NEXT.md
@@ -30,6 +30,7 @@ person cutting the release has any reason to re-read.
- [Ready for typed gRPC stubs (#61)](#ready-for-typed-grpc-stubs-issue-61)
- [Environment variables override the config file (#19)](#environment-variables-override-the-config-file-issue-19)
- [Detecting open-ended activations (#26)](#detecting-open-ended-activations-issue-26)
+- [Ingesting data (#17)](#ingesting-data-issue-17)
- [Cutting the release](#cutting-the-release)
---
@@ -140,6 +141,61 @@ not have its id.
See [#26](https://github.com/osprey-dcs/dp-python-lib/issues/26).
+## Ingesting data (Issue #17)
+
+The library can now put data into MLDP, not just read it. `client.ingestion_client`, which
+previously offered only `register_provider()`, now covers data ingestion and request status
+(`subscribeData()` is not wrapped yet):
+
+- **`ingest_data()`** sends one request. **`ingest_data_stream()`** sends many on one call and
+ returns a single summary. **`iter_ingest_data_bidi_stream()`** yields each request's ack or
+ reject as it arrives. To stop a bidi stream early, close its iterator, most simply with
+ `contextlib.closing`: that cancels the call. A plain `break` does not, while the iterator is still
+ referenced, and gRPC keeps sending the remaining requests in the background.
+- **`await_request_statuses()`** and **`query_request_status()`** report whether an ingestion
+ actually landed. This matters because **an ack means only that a request passed validation.**
+ The server queues the data and writes it afterwards, so a request can be acked and still fail.
+ An unknown provider id fails that way, and so does ingesting the same PV with the same first
+ timestamp twice. The request-status document, written once the data is stored, is the only
+ confirmation. Capture the time before sending and pass it as `since`.
+- **`IngestDataRequestParams`** takes a `common.DataFrame` built with the existing `data_frame`
+ builders, or with `data_frame_from_pandas()`. Its request id defaults to a generated uuid, since
+ the server does not enforce unique ids.
+- **`data_frame.split_data_frame()`** cuts a large frame into chunks that fit the server's inbound
+ message limit (4,096,000 bytes by default, about 500,000 doubles) and its one-day bucket span.
+ **`chunked_request_params()`** turns the chunks into requests with correlated ids, lazily, ready
+ for `ingest_data_stream()`. No limit is assumed: pass `SERVER_DEFAULT_MAX_MESSAGE_BYTES` to
+ target a default server.
+- **New column builders** for the kinds that had none: `double_array_column()` and its float, int32,
+ int64, and bool siblings (samples as nested lists or NumPy arrays), `image_column()`,
+ `struct_column()`, and `serialized_column()`.
+- **`RegisterProviderApiResult` gains `provider_id` and `is_new_provider`**, so the id no longer
+ has to be dug out of `.response.registrationResult`.
+- **`from_epoch_nanos()`**, the inverse of `to_epoch_nanos()`, is exported.
+
+Two behaviors worth knowing:
+
+- **A rejected request inside `ingest_data_stream()` does not raise.** The result has `is_error`
+ set and lists `rejected_request_ids`; the other requests were accepted.
+- **If your own request generator raises during a stream, you get your exception back**, not
+ gRPC's generic "Exception iterating requests!". Requests sent before it may or may not have
+ reached the server, so check their status.
+
+**Behavior change: `data_frame()` rejects hand-built columns it used to pass.** If you build
+column messages from the protobuf types directly and pass them to `data_frame()`, it now raises
+`ValueError` for an `EnumColumn` without an `enumId`, an array column with more than three
+dimensions, an `ImageColumn` without a complete image descriptor, a `StructColumn` without a
+`schemaId`, and a `SerializedDataColumn` without an `encoding`. The server rejects all of these
+anyway, so nothing that used to be accepted end to end is lost. `enum_column()` likewise now
+rejects a whitespace-only `enum_id`.
+
+Non-scalar columns (arrays, images, structs) can be ingested, but not yet read back through the
+query API, which returns scalar columns only; reading them back is
+[#16](https://github.com/osprey-dcs/dp-python-lib/issues/16).
+
+See [#17](https://github.com/osprey-dcs/dp-python-lib/issues/17) and the Ingestion API section of
+[`CLAUDE.md`](https://github.com/osprey-dcs/dp-python-lib/blob/main/CLAUDE.md).
+
## Installing
```bash
diff --git a/plan/tickets/17/plan.md b/plan/tickets/17/plan.md
index e01db36..421debb 100644
--- a/plan/tickets/17/plan.md
+++ b/plan/tickets/17/plan.md
@@ -237,6 +237,18 @@ All dp-service citations are `origin/main` @ `7e8b2e6`, paths relative to
`RpcError`) instead of returning a gRPC error. Its message says how many requests had already been sent,
because those were ingested (T4). Unit-tested against an in-process grpcio server, since a mocked stub
would not reproduce grpcio's swallowing.
+ - *Corrected in implementation (2026-09-30).* Two details above were wrong. (1) "Already sent" does not
+ mean "ingested", or even "received": the in-process tests show grpcio's cancellation racing the sends, so
+ the server holds some prefix of the requests handed over, possibly none. The note therefore says "any that
+ reached the server are ingested -- check request status" (T10's "have already been ingested" is corrected
+ the same way). (2) The count cannot go in the exception's *message* without changing its type or mutating
+ its args, so it travels as a PEP 678 note (Python 3.11+) and is always logged.
+ - *Added in implementation.* Closing the iterator `iter_ingest_data_bidi_stream()` returns cancels the call.
+ Without that, grpcio keeps pulling and sending the caller's requests on its own thread after the caller has
+ walked away; an in-process test showed all 1,000 queued requests drained into the server. The cancel needs an
+ explicit close: a `break` out of a loop over a generator that is still referenced leaves it suspended until
+ garbage collection, so the documented way to stop early is `with contextlib.closing(...)` (corrected in PR
+ review, 2026-09-30; the first draft said abandoning the iterator was enough).
- **D6 — `queryRequestStatus` gets a criterion helper and a poller.**
- `RequestStatusQuery` (`RS`): `provider_id(id)`, `provider_name(name)`, `request_id(id)`,
@@ -337,8 +349,10 @@ All dp-service citations are `origin/main` @ `7e8b2e6`, paths relative to
**`src/dp_python_lib/client/data_frame.py`**
- Extract `validate_data_frame(frame)` from `data_frame()`: axis via `timestamp_count()`, at least one
- column, then `_check_column()` on every column across all 16 arms. `data_frame()` calls it after
- assembly.
+ column, then `_check_column()` on every column across all 16 arms. *(As implemented, `data_frame()` does
+ not call it after assembly: both share one column-list check, which `data_frame()` runs on the caller's list
+ before routing, so an unsupported type is caught before it must be routed and messages keep the caller's
+ index.)*
- Extend `_check_column()` with T7's rules: enum `enumId` non-blank; array dims count 1–3; image
descriptor present with positive width/height/channels and non-blank encoding; struct `schemaId`
non-blank; serialized `encoding` non-blank.
diff --git a/src/dp_python_lib/client/__init__.py b/src/dp_python_lib/client/__init__.py
index 5f0b8b9..b2a487d 100644
--- a/src/dp_python_lib/client/__init__.py
+++ b/src/dp_python_lib/client/__init__.py
@@ -16,19 +16,29 @@
# `import dp_python_lib.client.data_frame` would hand back the function instead of the module -- breaking the
# documented `from dp_python_lib.client import data_frame as dfb` usage. Reach it as dfb.data_frame(...).
from dp_python_lib.client.data_frame import (
+ bool_array_column,
bool_column,
calculations_source,
column_metadata,
data_column,
+ double_array_column,
double_column,
enum_column,
+ float_array_column,
float_column,
+ image_column,
+ int32_array_column,
int32_column,
+ int64_array_column,
int64_column,
provenance,
pv_source,
+ serialized_column,
+ split_data_frame,
string_column,
+ struct_column,
timestamp_count,
+ validate_data_frame,
)
from dp_python_lib.client.dataset_client import (
DataSetClient,
@@ -48,9 +58,16 @@
calculations_spec,
)
from dp_python_lib.client.ingestion_client import (
+ IngestDataApiResult,
+ IngestDataRequestParams,
+ IngestDataStreamApiResult,
IngestionClient,
+ IngestionRequestStatus,
+ QueryRequestStatusApiResult,
RegisterProviderApiResult,
RegisterProviderRequestParams,
+ RequestStatusQuery,
+ chunked_request_params,
)
from dp_python_lib.client.machine_config_client import (
ConfigurationActivationQuery,
@@ -101,7 +118,7 @@
timestamp_list,
)
from dp_python_lib.client.sample_status_conversions import SampleStatusRow
-from dp_python_lib.client.time_conversions import TimestampInput, to_epoch_nanos, to_timestamp
+from dp_python_lib.client.time_conversions import TimestampInput, from_epoch_nanos, to_epoch_nanos, to_timestamp
__all__ = [
"AnnotationClient",
@@ -129,7 +146,11 @@
"GetConfigurationApiResult",
"GetDataSetApiResult",
"GetPvMetadataApiResult",
+ "IngestDataApiResult",
+ "IngestDataRequestParams",
+ "IngestDataStreamApiResult",
"IngestionClient",
+ "IngestionRequestStatus",
"MachineConfigClient",
"MldpClient",
"PvMetadataClient",
@@ -142,11 +163,13 @@
"QueryDataSetsApiResult",
"QueryParams",
"QueryPvMetadataApiResult",
+ "QueryRequestStatusApiResult",
"QuerySampleStatusesApiResult",
"QuerySampleStatusesRequestParams",
"QuerySamplesApiResult",
"RegisterProviderApiResult",
"RegisterProviderRequestParams",
+ "RequestStatusQuery",
"SampleStatusClient",
"SampleStatusColumn",
"SampleStatusFilter",
@@ -167,24 +190,36 @@
"TimestampInput",
"activation_end_time",
"activation_is_open",
+ "bool_array_column",
"bool_column",
"calculations",
"calculations_source",
"calculations_spec",
+ "chunked_request_params",
"column_metadata",
"data_block",
"data_column",
+ "double_array_column",
"double_column",
"enum_column",
+ "float_array_column",
"float_column",
+ "from_epoch_nanos",
+ "image_column",
+ "int32_array_column",
"int32_column",
+ "int64_array_column",
"int64_column",
"provenance",
"pv_source",
"sampling_clock",
+ "serialized_column",
+ "split_data_frame",
"string_column",
+ "struct_column",
"timestamp_count",
"timestamp_list",
"to_epoch_nanos",
"to_timestamp",
+ "validate_data_frame",
]
diff --git a/src/dp_python_lib/client/data_frame.py b/src/dp_python_lib/client/data_frame.py
index 5d9ae7c..8edc114 100644
--- a/src/dp_python_lib/client/data_frame.py
+++ b/src/dp_python_lib/client/data_frame.py
@@ -1,11 +1,11 @@
"""
-Builders for common.DataFrame -- the shared time-series payload shape (issue #6, Phase 2).
+Builders for common.DataFrame -- the shared time-series payload shape (issue #6, Phase 2; extended by issue #17).
A DataFrame is one time axis plus a set of columns sampled on it. It is the payload of an annotation's
Calculations, and it is also ingestion's `ingestionDataFrame`: the same message, so this module is the substrate
-issue #17 extends rather than a parallel one to fork.
+issue #17 extended rather than a parallel one to fork.
-Design decisions (see plan/tickets/6/plan.md, D6):
+Design decisions (see plan/tickets/6/plan.md, D6, and plan/tickets/17/plan.md, D1/D7/D8):
- The time axis builders sampling_clock() / timestamp_list() live here now that a second caller exists; they are
re-exported from sample_status_client so existing imports keep working.
- The axis carries its own sample count and every column is validated against it. That restates a number the
@@ -13,21 +13,41 @@
from the columns would fork the axis API by caller, since sample_status_client genuinely knows its count
independently of any column. SampleStatusFrame set the same precedent -- explicit axis, per-column validation.
- Validation mirrors the server's SHAPE rules only (non-blank names, non-empty values, count match, column-name
- uniqueness across types), so an error names the offending frame or column instead of bouncing the whole batch.
- The server's resource caps -- 256-char strings, 10M-element arrays, 50 MB images, 1 MB structs -- stay
- server-side: they are deployment policy, and duplicating numbers that can change is how clients drift.
- - Array, image, struct, and serialized column builders are deliberately absent. Their ergonomics (dims, image
- descriptors, schema ids) are #17's to design once with ingestion data in hand. Meanwhile they are reachable by
- building the proto directly and passing it to data_frame(), which accepts pre-built column messages alongside
- the ones its builders return.
+ uniqueness across types, and each column kind's structural fields: an enum's enumId, an array's 1-3 dims, an
+ image's descriptor, a struct's schemaId, a serialized column's encoding), so an error names the offending frame
+ or column instead of bouncing the whole batch. The structural rules live in _check_column(), not only in the
+ builders, because a hand-built column bypasses every builder; validate_data_frame() applies the same checks to
+ a whole frame, however it was made. The server's resource caps -- 256-char strings, 10M-element arrays,
+ 50 MB images, 1 MB structs -- stay server-side: they are deployment policy, and duplicating numbers that can
+ change is how clients drift.
+ - Two rules are stricter than the server's, deliberately: blank (whitespace-only) names and duplicate timestamps
+ are rejected, though ingestion accepts both. Anything data_frame() accepts must be readable back, the read
+ path rejects duplicate timestamps, and sample status matching assumes one sample per instant.
+ - split_data_frame() is the exception to "no deployment caps here": its whole purpose is to fit the server's
+ inbound message limit and bucket span cap, which a large frame hits at once (a 4 MB message holds roughly 500k
+ doubles; a day of one 10 kHz PV is 864M). So the caller names each limit explicitly, and
+ SERVER_DEFAULT_MAX_MESSAGE_BYTES is exported as a documented reference value, not applied as a default.
"""
-from collections.abc import Sequence
+from collections.abc import Iterator, Sequence
from numbers import Integral, Real
from typing import Any
-from dp_python_lib.client.time_conversions import TimestampInput, to_timestamp
-from dp_python_lib.grpc import common_pb2
+from dp_python_lib.client.time_conversions import TimestampInput, from_epoch_nanos, to_epoch_nanos, to_timestamp
+from dp_python_lib.grpc import common_pb2, ingestion_pb2
+
+# dp-service's default inbound gRPC message limit (common/server/GrpcServerBase.java, configurable as
+# GrpcServer.incomingMessageSizeLimitBytes). A documented reference for split_data_frame(max_bytes=...), not a
+# default: a deployment can change it, and a chunker that silently assumed it would drift with it.
+SERVER_DEFAULT_MAX_MESSAGE_BYTES = 4_096_000
+
+# The longest providerId / clientRequestId split_data_frame() budgets for. The chunker never sees the ids its
+# chunks will be sent with, so it reserves room for two ids of this many characters at the UTF-8 worst case of
+# 4 bytes each; IngestDataRequestParams rejects anything longer, so no request it builds can exceed the budget.
+MAX_BUDGETED_ID_CHARS = 256
+
+# Array columns carry 1-3 dimensions (the server rejects any other count).
+_MAX_ARRAY_DIMS = 3
# Each typed column message, paired with the DataFrame field it belongs in. data_frame() routes by exact type, so
# a pre-built column of any supported kind lands in the right repeated field without the caller naming it.
@@ -466,10 +486,10 @@ def enum_column(
:param enum_id: Id of the enumeration defining what the codes mean.
:param metadata: Optional per-column tags, attributes, and provenance (see column_metadata()).
:return: A common.EnumColumn.
- :raises ValueError: if name, values, or enum_id is empty.
+ :raises ValueError: if name or enum_id is blank, or values is empty.
"""
- if not enum_id:
- raise ValueError("enum_column() requires a non-empty enum_id")
+ if not enum_id or not enum_id.strip():
+ raise ValueError(f"enum_column() requires a non-blank enum_id, got {enum_id!r}")
column = _build_scalar_column(common_pb2.EnumColumn, "enum_column", name, values, metadata)
column.enumId = enum_id
return column
@@ -558,6 +578,346 @@ def data_column(
return column
+# ----------------------------------------------------------------------
+# array, image, struct, and serialized columns
+# ----------------------------------------------------------------------
+
+
+def _require_non_blank(value: str, builder_name: str, what: str) -> None:
+ """
+ Raises unless a string argument has content; whitespace-only counts as blank, as it does for column names.
+
+ :param value: The argument to check.
+ :param builder_name: The public builder's name, used in the message.
+ :param what: The argument's name, used in the message.
+ :raises ValueError: if value is empty or whitespace-only.
+ """
+ if not value or not value.strip():
+ raise ValueError(f"{builder_name}() requires a non-blank {what}, got {value!r}")
+
+
+def _sample_shape_and_values(sample: Any) -> tuple[tuple[int, ...], list]:
+ """
+ Returns one array sample's shape and its values flattened in row-major (C) order.
+
+ A sample is either anything with `.shape` and `.ravel()` -- a NumPy array, duck-typed so NumPy stays an
+ optional dependency -- or a nested sequence, whose shape is inferred from the nesting and must be rectangular.
+ Row-major is the order data_frame_conversions assumes when it slices a column back into samples.
+
+ :param sample: One sample of an array column.
+ :return: (shape, flat values); a scalar sample has the empty shape ().
+ :raises ValueError: if a nested sequence is ragged.
+ """
+ if hasattr(sample, "shape") and hasattr(sample, "ravel"):
+ # tolist() hands back Python scalars, which every repeated proto field accepts; NumPy scalars are not
+ # uniformly accepted (np.bool_ in a bool field, for one).
+ return tuple(int(extent) for extent in sample.shape), sample.ravel().tolist()
+ if isinstance(sample, (str, bytes)) or not isinstance(sample, Sequence):
+ return (), [sample]
+ if len(sample) == 0:
+ return (0,), []
+ shapes_and_values = [_sample_shape_and_values(element) for element in sample]
+ inner_shape = shapes_and_values[0][0]
+ flat: list = []
+ for position, (shape, values) in enumerate(shapes_and_values):
+ if shape != inner_shape:
+ raise ValueError(
+ f"element {position} has shape {list(shape)} but element 0 has shape {list(inner_shape)}; "
+ f"an array sample must be rectangular"
+ )
+ flat.extend(values)
+ return (len(sample), *inner_shape), flat
+
+
+def _build_array_column(
+ column_class: Any,
+ builder_name: str,
+ name: str,
+ samples: Sequence[Any],
+ dims: Sequence[int] | None,
+ metadata: common_pb2.ColumnMetadata | None,
+) -> Any:
+ """
+ Builds one array column: every sample the same shape, stored flat as samples x prod(dims) values.
+
+ :param column_class: The array column message class to construct.
+ :param builder_name: The public builder's name, used in error messages.
+ :param name: The column's name.
+ :param samples: One array per sample (nested sequences or NumPy arrays).
+ :param dims: Explicit per-sample dimensions, or None to use the samples' own shape.
+ :param metadata: Optional per-column metadata.
+ :return: The constructed array column message.
+ :raises ValueError: if a shape rule is violated; the message names the column and, where there is one, the
+ offending sample.
+ """
+ _require_non_blank(name, builder_name, "name")
+ # len(), not truthiness: a NumPy array of samples has no single truth value.
+ if len(samples) == 0:
+ raise ValueError(f"{builder_name}() requires at least one sample for column '{name}'")
+
+ flat: list = []
+ first_shape: tuple[int, ...] | None = None
+ for index, sample in enumerate(samples):
+ try:
+ shape, values = _sample_shape_and_values(sample)
+ except ValueError as e:
+ raise ValueError(f"{builder_name}() column '{name}' sample {index}: {e}") from e
+ if first_shape is None:
+ first_shape = shape
+ elif shape != first_shape:
+ raise ValueError(
+ f"{builder_name}() column '{name}' sample {index} has shape {list(shape)} but sample 0 has shape "
+ f"{list(first_shape)}; every sample in an array column must have the same shape"
+ )
+ flat.extend(values)
+ assert first_shape is not None # samples is non-empty
+
+ if len(first_shape) == 0:
+ raise ValueError(
+ f"{builder_name}() column '{name}' has scalar samples; each sample must be an array (use the scalar "
+ f"column builders for one value per sample)"
+ )
+
+ if dims is None:
+ resolved = list(first_shape)
+ else:
+ resolved = [int(dim) for dim in dims]
+ per_sample = 1
+ for dim in resolved:
+ per_sample *= dim
+ sample_size = len(flat) // len(samples)
+ # Explicit dims either restate the samples' own shape or give a flat sample its multidimensional shape.
+ # Anything else would silently reinterpret the data's layout (a 2x3 sample read as 3x2).
+ if not (list(first_shape) == resolved or (len(first_shape) == 1 and sample_size == per_sample)):
+ raise ValueError(
+ f"{builder_name}() column '{name}' samples have shape {list(first_shape)}, which dims {resolved} "
+ f"does not describe; pass dims only to restate that shape or to shape flat samples of "
+ f"{per_sample} values"
+ )
+
+ if not 1 <= len(resolved) <= _MAX_ARRAY_DIMS:
+ raise ValueError(
+ f"{builder_name}() column '{name}' has {len(resolved)} dimensions {resolved}; array columns support "
+ f"1 to {_MAX_ARRAY_DIMS}"
+ )
+ if any(dim <= 0 for dim in resolved):
+ raise ValueError(f"{builder_name}() column '{name}' has dimensions {resolved}; every dimension must be > 0")
+
+ column = column_class()
+ column.name = name
+ column.dimensions.dims.extend(resolved)
+ try:
+ column.values.extend(flat)
+ except (TypeError, ValueError) as e:
+ # protobuf raises TypeError for a wrong type and ValueError for an out-of-range integer; either way the
+ # message should name the column.
+ raise ValueError(f"{builder_name}() column '{name}' has a value its type cannot hold: {e}") from e
+ if metadata is not None:
+ column.metadata.CopyFrom(metadata)
+ return column
+
+
+def double_array_column(
+ name: str,
+ samples: Sequence[Any],
+ dims: Sequence[int] | None = None,
+ metadata: common_pb2.ColumnMetadata | None = None,
+) -> common_pb2.DoubleArrayColumn:
+ """
+ Builds a DoubleArrayColumn (float64): one fixed-shape array per sample, such as a waveform or a 2-D map.
+
+ Each sample is a nested sequence (its shape inferred, and required to be rectangular) or a NumPy array; a 2-D
+ NumPy array works as `samples` directly, one row per sample. Every sample must have the same shape, with 1 to
+ 3 dimensions. Values are stored flat in row-major order. Pass `dims` when the samples are already flat but
+ mean a multidimensional shape -- [64, 64] for 4096-value samples -- or to restate their shape; dims that would
+ reinterpret a multidimensional sample's layout are rejected.
+
+ :param name: The column's name, unique within its frame.
+ :param samples: One array per sample on the frame's time axis.
+ :param dims: Optional per-sample dimensions (see above).
+ :param metadata: Optional per-column tags, attributes, and provenance (see column_metadata()).
+ :return: A common.DoubleArrayColumn.
+ :raises ValueError: if name is blank, samples is empty, a sample is scalar or ragged, samples differ in shape,
+ dims do not describe the samples, or there are not 1 to 3 dimensions each > 0.
+ """
+ return _build_array_column(common_pb2.DoubleArrayColumn, "double_array_column", name, samples, dims, metadata)
+
+
+def float_array_column(
+ name: str,
+ samples: Sequence[Any],
+ dims: Sequence[int] | None = None,
+ metadata: common_pb2.ColumnMetadata | None = None,
+) -> common_pb2.FloatArrayColumn:
+ """
+ Builds a FloatArrayColumn (float32). Sample and dims rules are those of double_array_column().
+
+ :param name: The column's name, unique within its frame.
+ :param samples: One array per sample on the frame's time axis.
+ :param dims: Optional per-sample dimensions.
+ :param metadata: Optional per-column tags, attributes, and provenance (see column_metadata()).
+ :return: A common.FloatArrayColumn.
+ :raises ValueError: as for double_array_column().
+ """
+ return _build_array_column(common_pb2.FloatArrayColumn, "float_array_column", name, samples, dims, metadata)
+
+
+def int32_array_column(
+ name: str,
+ samples: Sequence[Any],
+ dims: Sequence[int] | None = None,
+ metadata: common_pb2.ColumnMetadata | None = None,
+) -> common_pb2.Int32ArrayColumn:
+ """
+ Builds an Int32ArrayColumn. Sample and dims rules are those of double_array_column().
+
+ :param name: The column's name, unique within its frame.
+ :param samples: One array per sample on the frame's time axis.
+ :param dims: Optional per-sample dimensions.
+ :param metadata: Optional per-column tags, attributes, and provenance (see column_metadata()).
+ :return: A common.Int32ArrayColumn.
+ :raises ValueError: as for double_array_column(), or if a value is not an integer in range.
+ """
+ return _build_array_column(common_pb2.Int32ArrayColumn, "int32_array_column", name, samples, dims, metadata)
+
+
+def int64_array_column(
+ name: str,
+ samples: Sequence[Any],
+ dims: Sequence[int] | None = None,
+ metadata: common_pb2.ColumnMetadata | None = None,
+) -> common_pb2.Int64ArrayColumn:
+ """
+ Builds an Int64ArrayColumn. Sample and dims rules are those of double_array_column().
+
+ :param name: The column's name, unique within its frame.
+ :param samples: One array per sample on the frame's time axis.
+ :param dims: Optional per-sample dimensions.
+ :param metadata: Optional per-column tags, attributes, and provenance (see column_metadata()).
+ :return: A common.Int64ArrayColumn.
+ :raises ValueError: as for double_array_column(), or if a value is not an integer in range.
+ """
+ return _build_array_column(common_pb2.Int64ArrayColumn, "int64_array_column", name, samples, dims, metadata)
+
+
+def bool_array_column(
+ name: str,
+ samples: Sequence[Any],
+ dims: Sequence[int] | None = None,
+ metadata: common_pb2.ColumnMetadata | None = None,
+) -> common_pb2.BoolArrayColumn:
+ """
+ Builds a BoolArrayColumn. Sample and dims rules are those of double_array_column().
+
+ :param name: The column's name, unique within its frame.
+ :param samples: One array per sample on the frame's time axis.
+ :param dims: Optional per-sample dimensions.
+ :param metadata: Optional per-column tags, attributes, and provenance (see column_metadata()).
+ :return: A common.BoolArrayColumn.
+ :raises ValueError: as for double_array_column().
+ """
+ return _build_array_column(common_pb2.BoolArrayColumn, "bool_array_column", name, samples, dims, metadata)
+
+
+def image_column(
+ name: str,
+ images: Sequence[bytes],
+ width: int,
+ height: int,
+ channels: int,
+ encoding: str,
+ metadata: common_pb2.ColumnMetadata | None = None,
+) -> common_pb2.ImageColumn:
+ """
+ Builds an ImageColumn: one encoded image per sample, all described by a single ImageDescriptor.
+
+ The descriptor is per column, as the proto has it, so every image in the column shares its width, height,
+ channel count, and encoding. The encoding is a producer/consumer contract ("png", "raw-u16le", ...) that
+ MLDP stores without interpreting.
+
+ :param name: The column's name, unique within its frame.
+ :param images: One encoded image per sample on the frame's time axis.
+ :param width: Image width in pixels. Must be > 0.
+ :param height: Image height in pixels. Must be > 0.
+ :param channels: Channels per pixel, e.g. 1 (mono) or 3 (RGB). Must be > 0.
+ :param encoding: How each image's bytes are encoded. Must be non-blank.
+ :param metadata: Optional per-column tags, attributes, and provenance (see column_metadata()).
+ :return: A common.ImageColumn.
+ :raises ValueError: if name or encoding is blank, images is empty, or a descriptor dimension is not positive.
+ """
+ _require_non_blank(name, "image_column", "name")
+ if len(images) == 0:
+ raise ValueError(f"image_column() requires at least one image for column '{name}'")
+ for label, value in (("width", width), ("height", height), ("channels", channels)):
+ if value <= 0:
+ raise ValueError(f"image_column() requires {label} > 0 for column '{name}', got {value}")
+ _require_non_blank(encoding, "image_column", "encoding")
+
+ column = common_pb2.ImageColumn()
+ column.name = name
+ column.imageDescriptor.width = width
+ column.imageDescriptor.height = height
+ column.imageDescriptor.channels = channels
+ column.imageDescriptor.encoding = encoding
+ column.images.extend(images)
+ if metadata is not None:
+ column.metadata.CopyFrom(metadata)
+ return column
+
+
+def struct_column(
+ name: str,
+ values: Sequence[bytes],
+ schema_id: str,
+ metadata: common_pb2.ColumnMetadata | None = None,
+) -> common_pb2.StructColumn:
+ """
+ Builds a StructColumn: one serialized structure per sample, all following one schema.
+
+ :param name: The column's name, unique within its frame.
+ :param values: One serialized structure per sample on the frame's time axis.
+ :param schema_id: Id of the schema the structures follow, e.g. "beam_position:v3". Must be non-blank.
+ :param metadata: Optional per-column tags, attributes, and provenance (see column_metadata()).
+ :return: A common.StructColumn.
+ :raises ValueError: if name or schema_id is blank, or values is empty.
+ """
+ _require_non_blank(schema_id, "struct_column", "schema_id")
+ column = _build_scalar_column(common_pb2.StructColumn, "struct_column", name, list(values), metadata)
+ column.schemaId = schema_id
+ return column
+
+
+def serialized_column(
+ name: str,
+ payload: bytes,
+ encoding: str,
+ metadata: common_pb2.ColumnMetadata | None = None,
+) -> common_pb2.SerializedDataColumn:
+ """
+ Builds a SerializedDataColumn: a whole column's values as one opaque byte string.
+
+ The payload is opaque to MLDP and to this library, so it carries no sample count and is not checked against
+ the time axis -- the server checks only its name and encoding. For the same reason split_data_frame() cannot
+ divide a frame containing one.
+
+ :param name: The column's name, unique within its frame.
+ :param payload: The encoded column.
+ :param encoding: How the payload is encoded, e.g. "proto:Image". Must be non-blank.
+ :param metadata: Optional per-column tags, attributes, and provenance (see column_metadata()).
+ :return: A common.SerializedDataColumn.
+ :raises ValueError: if name or encoding is blank.
+ """
+ _require_non_blank(name, "serialized_column", "name")
+ _require_non_blank(encoding, "serialized_column", "encoding")
+ column = common_pb2.SerializedDataColumn()
+ column.name = name
+ column.encoding = encoding
+ column.payload = payload
+ if metadata is not None:
+ column.metadata.CopyFrom(metadata)
+ return column
+
+
# ----------------------------------------------------------------------
# frame assembly
# ----------------------------------------------------------------------
@@ -619,7 +979,7 @@ def _column_sample_count(column: Any) -> int | None:
return len(column.values)
-def _check_column(column: Any, index: int, expected_count: int, seen_names: set[str]) -> None:
+def _check_column(column: Any, index: int, expected_count: int, seen_names: set[str], caller: str) -> None:
"""
Applies the server's per-column shape rules to one column, with a message naming the column.
@@ -627,56 +987,128 @@ def _check_column(column: Any, index: int, expected_count: int, seen_names: set[
:param index: The column's position in the caller's list, for messages about an unnamed column.
:param expected_count: The frame's sample count, which every counted column must match.
:param seen_names: Names already used in this frame; mutated to record this column's name.
- :raises ValueError: if the column is of an unsupported type, unnamed, empty, duplicate-named, or
- count-mismatched.
+ :param caller: The public entry point, as it should appear at the start of an error message.
+ :raises ValueError: if the column is of an unsupported type, unnamed, empty, duplicate-named, count-mismatched,
+ or missing a structural field its kind requires (enumId, 1-3 array dims, an image descriptor, schemaId,
+ or a serialized column's encoding).
"""
if type(column) not in _COLUMN_FIELD_BY_TYPE:
raise ValueError(
- f"data_frame() received an unsupported column type at index {index}: {type(column).__name__}. "
+ f"{caller} received an unsupported column type at index {index}: {type(column).__name__}. "
f"Use one of the column builders, or pass a pre-built typed column message."
)
name = column.name
if not name or not name.strip():
# Blank, not merely empty: the rule is a non-blank name, and a whitespace-only one is not a name.
- raise ValueError(
- f"data_frame() requires a non-blank name for every column; column at index {index} has {name!r}"
- )
+ raise ValueError(f"{caller} requires a non-blank name for every column; column at index {index} has {name!r}")
if name in seen_names:
raise ValueError(
- f"data_frame() requires unique column names within a frame; '{name}' appears more than once "
+ f"{caller} requires unique column names within a frame; '{name}' appears more than once "
f"(names must be unique across ALL column types, not just within one type)"
)
seen_names.add(name)
+ # Each kind's structural fields. The builders enforce these too, but a hand-built column bypasses every
+ # builder, and the server rejects all of them -- so this is the check that holds for any column (#17, T7).
if isinstance(column, common_pb2.SerializedDataColumn):
- # Serialized payloads are opaque; the server checks their names only.
+ # Serialized payloads are opaque; the server checks only the name and the encoding, never a count.
+ if not column.encoding.strip():
+ raise ValueError(f"{caller} serialized column '{name}' requires a non-blank encoding")
return
+ if isinstance(column, common_pb2.EnumColumn) and not column.enumId.strip():
+ raise ValueError(f"{caller} enum column '{name}' requires a non-blank enumId")
+ if isinstance(column, common_pb2.StructColumn) and not column.schemaId.strip():
+ raise ValueError(f"{caller} struct column '{name}' requires a non-blank schemaId")
+ if isinstance(column, common_pb2.ImageColumn):
+ if not column.HasField("imageDescriptor"):
+ raise ValueError(f"{caller} image column '{name}' requires an imageDescriptor")
+ descriptor = column.imageDescriptor
+ for label, value in (
+ ("width", descriptor.width),
+ ("height", descriptor.height),
+ ("channels", descriptor.channels),
+ ):
+ if value <= 0:
+ raise ValueError(f"{caller} image column '{name}' requires imageDescriptor.{label} > 0")
+ if not descriptor.encoding.strip():
+ raise ValueError(f"{caller} image column '{name}' requires a non-blank imageDescriptor.encoding")
if isinstance(column, _ARRAY_COLUMN_TYPES):
+ dims = list(column.dimensions.dims)
+ if len(dims) > _MAX_ARRAY_DIMS:
+ raise ValueError(
+ f"{caller} array column '{name}' has {len(dims)} dimensions {dims}; array columns support "
+ f"1 to {_MAX_ARRAY_DIMS}"
+ )
product = _array_sample_size(column)
if product is None:
raise ValueError(
- f"data_frame() cannot determine the sample count of array column '{name}': its dimensions are "
+ f"{caller} cannot determine the sample count of array column '{name}': its dimensions are "
f"missing or zero. Set ArrayDimensions.dims so that values is samples x prod(dims)."
)
if len(column.values) % product != 0:
raise ValueError(
- f"data_frame() array column '{name}' has {len(column.values)} values, which is not a whole "
+ f"{caller} array column '{name}' has {len(column.values)} values, which is not a whole "
f"multiple of its per-sample size {product} (from dims {list(column.dimensions.dims)}); "
f"values must be samples x prod(dims)"
)
count = _column_sample_count(column)
if count == 0:
- raise ValueError(f"data_frame() requires a non-empty values list for column '{name}'")
+ raise ValueError(f"{caller} requires a non-empty values list for column '{name}'")
if count != expected_count:
raise ValueError(
- f"data_frame() column '{name}' has {count} values but the time axis has {expected_count} timestamps; "
+ f"{caller} column '{name}' has {count} values but the time axis has {expected_count} timestamps; "
f"every column must carry exactly one value per sample"
)
+def _check_columns(columns: Sequence[Any], expected_count: int, caller: str) -> None:
+ """
+ Applies _check_column() to every column of one frame, sharing a single set of names across them all.
+
+ :param columns: The frame's columns, in any mix of supported types.
+ :param expected_count: The frame's sample count.
+ :param caller: The public entry point, as it should appear at the start of an error message.
+ :raises ValueError: as for _check_column(), or if there are no columns.
+ """
+ if not columns:
+ raise ValueError(f"{caller} requires at least one column")
+ seen_names: set[str] = set()
+ for index, column in enumerate(columns):
+ _check_column(column, index, expected_count, seen_names, caller)
+
+
+def _frame_columns(frame: common_pb2.DataFrame) -> list[Any]:
+ """
+ Returns every column of a frame, across all of its repeated fields, serialized columns included.
+
+ :param frame: The frame to walk.
+ :return: The frame's column messages, grouped by type in _COLUMN_FIELD_BY_TYPE order.
+ """
+ return [column for field in _COLUMN_FIELD_BY_TYPE.values() for column in getattr(frame, field)]
+
+
+def validate_data_frame(frame: common_pb2.DataFrame, *, caller: str = "validate_data_frame()") -> None:
+ """
+ Applies data_frame()'s checks to an already-assembled frame, however it was made.
+
+ data_frame() validates as it assembles, but a frame can also come from data_frame_from_pandas(),
+ split_data_frame(), or a hand-built message, and ingestion re-validates whatever it is given here so every
+ path fails with the same client-side messages. A column's "index" in an error counts across the frame's
+ repeated fields in type order, since an assembled frame no longer has the caller's original ordering.
+
+ :param frame: The frame to check.
+ :param caller: The name an error message starts with. A function that validates a frame on its caller's
+ behalf passes its own name, so the message names what the caller actually called.
+ :raises ValueError: if the axis is empty or malformed, if the frame has no columns, or if any column violates a
+ shape rule (see data_frame()).
+ """
+ expected_count = timestamp_count(frame.dataTimestamps)
+ _check_columns(_frame_columns(frame), expected_count, caller)
+
+
def data_frame(
data_timestamps: common_pb2.DataTimestamps,
columns: list[Any],
@@ -685,13 +1117,15 @@ def data_frame(
Assembles a common.DataFrame from a time axis and a list of columns, routing each column into the repeated
field for its type.
- Accepts the columns this module builds and pre-built column messages alike, so the kinds without builders yet
- (arrays, images, structs, serialized payloads -- issue #17) can be passed straight through.
+ Accepts the columns this module builds and pre-built column messages alike.
Validates the server's shape rules client-side, so a mistake names the offending column instead of bouncing the
whole save: at least one column, a non-blank name and non-empty values for each, a value count matching the
- axis (for arrays, values / prod(dims)), and column names unique across ALL types within the frame. The
- server's size caps are not duplicated -- see the module docstring.
+ axis (for arrays, values / prod(dims)), column names unique across ALL types within the frame, and each kind's
+ structural fields -- an enum's enumId, an array's 1 to 3 dims, an image's descriptor (positive width, height,
+ and channels, and an encoding), a struct's schemaId, and a serialized column's encoding. The server's size
+ caps are not duplicated -- see the module docstring. validate_data_frame() applies the same checks to a frame
+ built any other way.
:param data_timestamps: The frame's time axis (see sampling_clock() / timestamp_list()).
:param columns: The frame's columns, in any mix of supported types.
@@ -702,14 +1136,253 @@ def data_frame(
if not columns:
raise ValueError("data_frame() requires at least one column")
- expected_count = timestamp_count(data_timestamps)
-
- seen_names: set[str] = set()
- for index, column in enumerate(columns):
- _check_column(column, index, expected_count, seen_names)
+ # Checked here, against the caller's list, rather than by validate_data_frame() after assembly: an unsupported
+ # type cannot be routed at all, and the caller's own index is the useful one in a message.
+ _check_columns(columns, timestamp_count(data_timestamps), "data_frame()")
frame = common_pb2.DataFrame()
frame.dataTimestamps.CopyFrom(data_timestamps)
for column in columns:
getattr(frame, _COLUMN_FIELD_BY_TYPE[type(column)]).append(column)
return frame
+
+
+# ----------------------------------------------------------------------
+# chunking
+# ----------------------------------------------------------------------
+
+
+def _varint_size(value: int) -> int:
+ """
+ Returns the number of bytes protobuf's varint encoding uses for a non-negative integer.
+
+ :param value: The integer to encode.
+ :return: Its encoded length, 1 to 10 bytes.
+ """
+ return max(1, (value.bit_length() + 6) // 7)
+
+
+# The serialized size of an IngestDataRequest's two id fields at their budgeted worst case: MAX_BUDGETED_ID_CHARS
+# characters that each take the maximum 4 UTF-8 bytes. Measured from the message rather than hard-coded, so it
+# stays right if the field numbers or wire format ever change.
+_WORST_CASE_ID = "\U0010ffff" * MAX_BUDGETED_ID_CHARS
+_BUDGETED_ID_FIELDS_BYTES = ingestion_pb2.IngestDataRequest(
+ providerId=_WORST_CASE_ID, clientRequestId=_WORST_CASE_ID
+).ByteSize()
+
+# The ingestionDataFrame field's tag. Field 3, length-delimited, encodes in one byte.
+_FRAME_FIELD_TAG_BYTES = 1
+
+
+def _budgeted_request_bytes(frame_bytes: int) -> int:
+ """
+ Returns the size of an IngestDataRequest carrying a frame of the given size and two worst-case ids.
+
+ This is what split_data_frame() compares against max_bytes: the whole message the server's inbound limit
+ applies to, not the frame alone. The envelope is small but not fixed -- the ids are the caller's strings, and
+ the frame's length prefix grows with the frame -- so it is computed rather than guessed.
+
+ :param frame_bytes: The frame's serialized size (DataFrame.ByteSize()).
+ :return: The request's serialized size, in bytes.
+ """
+ return _BUDGETED_ID_FIELDS_BYTES + _FRAME_FIELD_TAG_BYTES + _varint_size(frame_bytes) + frame_bytes
+
+
+def _sample_field(column: Any) -> str:
+ """
+ Returns the name of the repeated field holding a column's per-sample values.
+
+ :param column: Any countable column (not a SerializedDataColumn).
+ :return: "dataValues" for a DataColumn, "images" for an ImageColumn, otherwise "values".
+ """
+ if isinstance(column, common_pb2.DataColumn):
+ return "dataValues"
+ if isinstance(column, common_pb2.ImageColumn):
+ return "images"
+ return "values"
+
+
+def _fill_column_slice(target: Any, column: Any, start: int, end: int) -> None:
+ """
+ Fills an empty column with rows [start, end) of another column of the same type.
+
+ Every field other than the per-sample values -- name, metadata, enumId, schemaId, dimensions, image
+ descriptor -- is copied whole, generically, so a field added to a column message later is carried into every
+ chunk without a change here. Array values are sliced in prod(dims)-sized blocks, one block per row.
+
+ :param target: A new, empty column message of the same type as column.
+ :param column: The column to slice.
+ :param start: First row, inclusive.
+ :param end: Last row, exclusive.
+ """
+ sample_field = _sample_field(column)
+ for field, value in column.ListFields():
+ if field.name == sample_field:
+ continue
+ if field.message_type is not None and not field.is_repeated:
+ getattr(target, field.name).CopyFrom(value)
+ elif field.is_repeated:
+ getattr(target, field.name).extend(value)
+ else:
+ setattr(target, field.name, value)
+
+ block = 1
+ if isinstance(column, _ARRAY_COLUMN_TYPES):
+ # Validated before chunking starts, so the dims are present and positive.
+ block = _array_sample_size(column) or 1
+ getattr(target, sample_field).extend(getattr(column, sample_field)[start * block : end * block])
+
+
+def _slice_frame(frame: common_pb2.DataFrame, start: int, end: int) -> common_pb2.DataFrame:
+ """
+ Returns rows [start, end) of a frame as a new frame carrying every column.
+
+ A SamplingClock chunk keeps the period and gets start time startTime + start * periodNanos, computed in integer
+ nanoseconds: present-day epoch nanoseconds need about 61 bits, and a float would move the instant. A
+ TimestampList is sliced.
+
+ :param frame: A validated frame with no serialized columns.
+ :param start: First row, inclusive.
+ :param end: Last row, exclusive.
+ :return: The chunk.
+ """
+ chunk = common_pb2.DataFrame()
+ axis = frame.dataTimestamps
+ if axis.WhichOneof("value") == "samplingClock":
+ clock = chunk.dataTimestamps.samplingClock
+ first_nanos = to_epoch_nanos(axis.samplingClock.startTime) + start * axis.samplingClock.periodNanos
+ clock.startTime.CopyFrom(from_epoch_nanos(first_nanos))
+ clock.periodNanos = axis.samplingClock.periodNanos
+ clock.count = end - start
+ else:
+ chunk.dataTimestamps.timestampList.timestamps.extend(axis.timestampList.timestamps[start:end])
+
+ for field in _COLUMN_FIELD_BY_TYPE.values():
+ for column in getattr(frame, field):
+ _fill_column_slice(getattr(chunk, field).add(), column, start, end)
+ return chunk
+
+
+def split_data_frame(
+ frame: common_pb2.DataFrame,
+ *,
+ max_rows: int | None = None,
+ max_bytes: int | None = None,
+ max_span_nanos: int | None = None,
+) -> Iterator[common_pb2.DataFrame]:
+ """
+ Splits a frame along its time axis into chunks that each fit the given limits, lazily.
+
+ A large frame hits the server's limits at once: its inbound message cap (about 4 MB by default -- roughly 500k
+ doubles) and its bucket span cap (1 day by default). Over the message cap the call fails with a gRPC status,
+ and on a stream it kills every request after it. Each chunk carries every column, with its metadata, over a
+ contiguous run of rows; concatenating the chunks' rows reproduces the frame.
+
+ - max_rows caps the rows per chunk.
+ - max_bytes caps the size of the whole IngestDataRequest a chunk will travel in, not just the frame, with room
+ reserved for a providerId and clientRequestId of up to MAX_BUDGETED_ID_CHARS characters each. So a chunk
+ within the budget fits the message; pass SERVER_DEFAULT_MAX_MESSAGE_BYTES to target a default server.
+ Chunks are sized from the rows already measured, so they fit but are not guaranteed to be the largest that
+ would.
+ - max_span_nanos caps the time from a chunk's first timestamp to its last, as the server measures the bucket
+ span: a chunk may span exactly the limit. Pass 86_400 * NANOS_PER_SECOND to target a default server.
+
+ The server's limits are deployment settings, so none is assumed: at least one must be given.
+
+ The frame is validated (see validate_data_frame()) before the first chunk is produced, and the call raises
+ immediately rather than on first iteration. Chunks are built as they are consumed, so feeding them to a
+ streaming ingest never holds every chunk in memory at once.
+
+ :param frame: The frame to split.
+ :param max_rows: Maximum rows per chunk. Must be > 0.
+ :param max_bytes: Maximum size, in bytes, of the IngestDataRequest carrying each chunk. Must be > 0.
+ :param max_span_nanos: Maximum nanoseconds from a chunk's first timestamp to its last. Must be >= 0.
+ :return: An iterator over the chunks, in time order.
+ :raises ValueError: if no limit is given or a limit is out of range, if the frame is invalid, if it contains a
+ SerializedDataColumn (whose opaque payload cannot be divided), or -- during iteration -- if a single row
+ exceeds max_bytes on its own.
+ """
+ if max_rows is None and max_bytes is None and max_span_nanos is None:
+ raise ValueError("split_data_frame() requires at least one of max_rows, max_bytes, or max_span_nanos")
+ if max_rows is not None and max_rows <= 0:
+ raise ValueError(f"split_data_frame() requires max_rows > 0, got {max_rows}")
+ if max_bytes is not None and max_bytes <= 0:
+ raise ValueError(f"split_data_frame() requires max_bytes > 0, got {max_bytes}")
+ if max_span_nanos is not None and max_span_nanos < 0:
+ raise ValueError(f"split_data_frame() requires max_span_nanos >= 0, got {max_span_nanos}")
+
+ validate_data_frame(frame, caller="split_data_frame()")
+ if len(frame.serializedDataColumns) > 0:
+ names = [column.name for column in frame.serializedDataColumns]
+ raise ValueError(
+ f"split_data_frame() cannot split a frame with serialized columns {names}: a serialized payload is "
+ f"opaque and cannot be divided by row"
+ )
+
+ return _iter_chunks(frame, max_rows, max_bytes, max_span_nanos)
+
+
+def _iter_chunks(
+ frame: common_pb2.DataFrame,
+ max_rows: int | None,
+ max_bytes: int | None,
+ max_span_nanos: int | None,
+) -> Iterator[common_pb2.DataFrame]:
+ """
+ The generator behind split_data_frame(), which validates its arguments eagerly and then returns this.
+
+ :param frame: A validated frame with no serialized columns.
+ :param max_rows: Maximum rows per chunk, or None.
+ :param max_bytes: Maximum request bytes per chunk, or None.
+ :param max_span_nanos: Maximum first-to-last nanoseconds per chunk, or None.
+ :return: An iterator over the chunks.
+ :raises ValueError: if a single row exceeds max_bytes.
+ """
+ row_count = timestamp_count(frame.dataTimestamps)
+ axis = frame.dataTimestamps
+ clock_period = axis.samplingClock.periodNanos if axis.WhichOneof("value") == "samplingClock" else None
+ list_nanos = (
+ None if clock_period is not None else [to_epoch_nanos(entry) for entry in axis.timestampList.timestamps]
+ )
+
+ # Bytes per row, first estimated from the whole frame and then from each chunk as it is measured, so the next
+ # chunk's first guess tracks the data as it changes.
+ row_bytes_estimate = max(1, frame.ByteSize() // row_count) if max_bytes is not None else 1
+
+ start = 0
+ while start < row_count:
+ end = row_count
+ if max_rows is not None:
+ end = min(end, start + max_rows)
+ if max_span_nanos is not None:
+ if clock_period is not None:
+ end = min(end, start + max_span_nanos // clock_period + 1)
+ else:
+ assert list_nanos is not None
+ limit = list_nanos[start] + max_span_nanos
+ span_end = start + 1
+ while span_end < end and list_nanos[span_end] <= limit:
+ span_end += 1
+ end = span_end
+
+ if max_bytes is None:
+ yield _slice_frame(frame, start, end)
+ start = end
+ continue
+
+ end = min(end, start + max(1, max_bytes // row_bytes_estimate))
+ while True:
+ chunk = _slice_frame(frame, start, end)
+ size = _budgeted_request_bytes(chunk.ByteSize())
+ if size <= max_bytes:
+ break
+ if end - start == 1:
+ raise ValueError(
+ f"split_data_frame() row {start} alone makes a {size}-byte request, over max_bytes {max_bytes} "
+ f"(which includes room for {MAX_BUDGETED_ID_CHARS}-character ids); no split can fit it"
+ )
+ # Shrink in proportion to the overshoot; the loop repeats if the rows are uneven.
+ end = min(end - 1, start + max(1, (end - start) * max_bytes // size))
+ row_bytes_estimate = max(1, size // (end - start))
+ yield chunk
+ start = end
diff --git a/src/dp_python_lib/client/data_frame_conversions.py b/src/dp_python_lib/client/data_frame_conversions.py
index 50e875e..a268425 100644
--- a/src/dp_python_lib/client/data_frame_conversions.py
+++ b/src/dp_python_lib/client/data_frame_conversions.py
@@ -23,7 +23,7 @@
from dp_python_lib.client.query_conversions import data_value_to_python
from dp_python_lib.client.sample_status_conversions import expand_data_timestamps
-from dp_python_lib.client.time_conversions import to_epoch_nanos
+from dp_python_lib.client.time_conversions import from_epoch_nanos, to_epoch_nanos
from dp_python_lib.grpc import annotation_pb2, common_pb2
# The DataFrame fields holding typed scalar columns, in the proto's declaration order. Each carries `values`
@@ -545,18 +545,6 @@ def _timestamps_from_index(index: Any) -> common_pb2.DataTimestamps:
return timestamps
-def _timestamp_from_nanos(epoch_nanos: int) -> common_pb2.Timestamp:
- """
- Builds a common.Timestamp from integer epoch nanoseconds -- the inverse of to_epoch_nanos().
-
- :param epoch_nanos: Epoch nanoseconds.
- :return: The equivalent common.Timestamp.
- """
- timestamp = common_pb2.Timestamp()
- timestamp.epochSeconds, timestamp.nanoseconds = divmod(epoch_nanos, 1_000_000_000)
- return timestamp
-
-
def column_metadata_from_dict(summary: dict[str, Any] | None) -> common_pb2.ColumnMetadata | None:
"""
Rebuilds a ColumnMetadata from the dict column_metadata_dict() produced -- the inverse of that function.
@@ -604,8 +592,8 @@ def column_metadata_from_dict(summary: dict[str, Any] | None) -> common_pb2.Colu
time_range = entry.get("time_range")
if time_range is not None:
begin_nanos, end_nanos = time_range
- source.timeRange.beginTime.CopyFrom(_timestamp_from_nanos(begin_nanos))
- source.timeRange.endTime.CopyFrom(_timestamp_from_nanos(end_nanos))
+ source.timeRange.beginTime.CopyFrom(from_epoch_nanos(begin_nanos))
+ source.timeRange.endTime.CopyFrom(from_epoch_nanos(end_nanos))
return metadata
diff --git a/src/dp_python_lib/client/ingestion_client.py b/src/dp_python_lib/client/ingestion_client.py
index 801a5ca..cf58fcc 100644
--- a/src/dp_python_lib/client/ingestion_client.py
+++ b/src/dp_python_lib/client/ingestion_client.py
@@ -1,15 +1,144 @@
+"""
+Ingestion Service client: provider registration, data ingestion (unary, client-streaming, bidirectional), and the
+request-status query that says whether an ingestion actually landed.
+
+Plan: plan/tickets/17/plan.md. Server behaviors this module encodes, all verified against dp-service:
+
+ - **An ack means "passed validation", not "ingested".** The service validates a request, acks or rejects it, and
+ only then queues it for asynchronous handling. An unknown providerId is acked and then fails asynchronously, as
+ do bucket-building errors and Mongo insert failures -- including a re-ingest of the same PV with the same first
+ timestamp, which can leave a partial write behind an ERROR status. The request-status document, written after
+ the data, is the only confirmation: see query_request_status() and await_request_statuses().
+ - **IngestionRequestStatus's zero value is SUCCESS**, so an unset status field reads as success. Nothing here
+ ever defaults a status; compare against IngestionRequestStatus members, and treat a missing document as unknown.
+ - **ingestDataStream reports rejects on an error response**: if any request in the stream was rejected, the
+ single response is an ExceptionalResult, and its rejectedRequestIds are the only record of which. So
+ IngestDataStreamApiResult keeps the response on error, unlike every unary result.
+ - **grpcio hides exceptions raised by a request iterator**: the caller sees UNKNOWN "Exception iterating
+ requests!". The streaming methods re-raise the original exception instead (_RequestFeed). Requests handed to
+ gRPC before it may or may not have reached the server -- the cancellation races their delivery -- and any that
+ did are ingested and stay so. Request status is the only way to tell which.
+"""
+
+import contextlib
import logging
+import time
+import uuid
+from collections.abc import Callable, Generator, Iterable, Iterator, Sequence
+from datetime import timedelta
+from enum import IntEnum
import grpc
+from dp_python_lib.client.data_frame import (
+ MAX_BUDGETED_ID_CHARS,
+ SERVER_DEFAULT_MAX_MESSAGE_BYTES,
+ validate_data_frame,
+)
from dp_python_lib.client.result import ApiResultBase
from dp_python_lib.client.service_api_client_base import ServiceApiClientBase
+from dp_python_lib.client.time_conversions import (
+ NANOS_PER_SECOND,
+ TimestampInput,
+ from_epoch_nanos,
+ to_epoch_nanos,
+ to_timestamp,
+)
from dp_python_lib.grpc import common_pb2, ingestion_pb2, ingestion_pb2_grpc
+# The floor await_request_statuses() puts under its status query is backed off by this much, because the status
+# documents' createdAt comes from the server's clock and `since` from the caller's.
+REQUEST_STATUS_CLOCK_SKEW = timedelta(seconds=60)
+
+# Status-detail texts that identify a message-size violation, matched case-insensitively within a
+# RESOURCE_EXHAUSTED. RESOURCE_EXHAUSTED alone also covers quotas and other resource limits, where "split your
+# frame" would be the wrong advice. Neither text is a stable API, so both are kept here, with their sources.
+_SERVER_INBOUND_SIZE_TEXT = "grpc message exceeds maximum size" # grpc-java server, over its inbound limit
+_CLIENT_RECEIVE_SIZE_TEXT = "received message larger than max" # grpcio client, over its receive limit
+
+# TimeRangeCriterion.beginTime must be at least 1 s past the epoch, or the server rejects the criterion.
+_MIN_STATUS_QUERY_BEGIN_NANOS = NANOS_PER_SECOND
+
+
+def _resource_exhausted_details(e: grpc.RpcError) -> str | None:
+ """
+ Returns a RESOURCE_EXHAUSTED error's details, lowercased, or None for any other status.
+
+ A bare grpc.RpcError (as the unit-test mocks raise) has no usable code(), so this is defensive about it.
+
+ :param e: The caught error.
+ :return: The lowercased details for RESOURCE_EXHAUSTED, else None.
+ """
+ try:
+ if e.code() != grpc.StatusCode.RESOURCE_EXHAUSTED:
+ return None
+ return str(e.details() or "").lower()
+ except (AttributeError, TypeError):
+ return None
+
+
+def _ingest_size_hint(e: grpc.RpcError) -> str:
+ """
+ Advice to append to an ingest call's gRPC error when the server refused the message as too large.
+
+ :param e: The caught error.
+ :return: The hint, or "" when the error is not the server's inbound message limit.
+ """
+ details = _resource_exhausted_details(e)
+ if details is None or _SERVER_INBOUND_SIZE_TEXT not in details:
+ return ""
+ return (
+ f" (the request exceeds the server's inbound message limit, {SERVER_DEFAULT_MAX_MESSAGE_BYTES:,} bytes by "
+ "default; split the frame with data_frame.split_data_frame(frame, max_bytes=...) and ingest the chunks, "
+ "e.g. with ingest_data_stream())"
+ )
+
+
+def _status_query_size_hint(e: grpc.RpcError) -> str:
+ """
+ Advice to append to a request-status query's gRPC error when its response exceeded the client's receive limit.
+
+ :param e: The caught error.
+ :return: The hint, or "" when the error is not the client's receive limit.
+ """
+ details = _resource_exhausted_details(e)
+ if details is None or _CLIENT_RECEIVE_SIZE_TEXT not in details:
+ return ""
+ return (
+ " (the response exceeds this client's receive limit: queryRequestStatus returns every match in one message, "
+ "with no paging yet, so narrow the query -- a later time_range begin, or a status criterion)"
+ )
+
+
+def _require_id(value: str, what: str) -> None:
+ """
+ Raises unless an id is non-blank and within the length split_data_frame() budgets for.
+
+ :param value: The id to check.
+ :param what: Its parameter name, for the message.
+ :raises ValueError: if value is blank or longer than MAX_BUDGETED_ID_CHARS characters.
+ """
+ if not value or not value.strip():
+ raise ValueError(f"{what} must be non-blank, got {value!r}")
+ if len(value) > MAX_BUDGETED_ID_CHARS:
+ raise ValueError(
+ f"{what} is {len(value)} characters; at most {MAX_BUDGETED_ID_CHARS} are allowed, the length "
+ f"split_data_frame() reserves room for in each request"
+ )
+
+
+# ----------------------------------------------------------------------
+# provider registration
+# ----------------------------------------------------------------------
+
class RegisterProviderRequestParams:
"""
Encapsulates client parameters for call to registerProvider() API method.
+
+ Registering an existing name is not an error: it returns that provider's id with is_new_provider False, and
+ updates it in place. The description is overwritten -- to "" if omitted -- but an empty tag list or attribute
+ map does not clear the stored ones.
"""
def __init__(
@@ -50,11 +179,410 @@ def __init__(
super().__init__(is_error, message)
self.response = response
+ @property
+ def provider_id(self) -> str | None:
+ """The provider's id, for use in every ingestion request; None on error."""
+ if self.response is None or not self.response.HasField("registrationResult"):
+ return None
+ return self.response.registrationResult.providerId
+
+ @property
+ def is_new_provider(self) -> bool | None:
+ """True if this call created the provider, False if it already existed; None on error."""
+ if self.response is None or not self.response.HasField("registrationResult"):
+ return None
+ return self.response.registrationResult.isNewProvider
+
+
+# ----------------------------------------------------------------------
+# data ingestion
+# ----------------------------------------------------------------------
+
+
+class IngestDataRequestParams:
+ """
+ Encapsulates client parameters for one ingestion request: a provider, a frame, and a request id.
+
+ The frame is a common.DataFrame from data_frame.data_frame(), data_frame_conversions.data_frame_from_pandas(),
+ data_frame.split_data_frame(), or built by hand. It is re-validated here with the same checks data_frame()
+ applies, so every path fails with the same client-side messages. Those checks are stricter than the server's
+ in two places -- whitespace-only column names and duplicate timestamps are rejected -- because the library must
+ be able to read back what it writes; a caller who needs either can build the request and call the stub directly.
+
+ Validation happens at construction, so a request built up front fails where it is made rather than
+ mid-stream. Params built lazily -- by chunked_request_params(), or any generator feeding a streaming ingest --
+ are constructed as the stream consumes them, so there a bad frame does fail mid-stream, after earlier requests
+ may already have been sent (see ingest_data_stream()).
+ """
+
+ def __init__(
+ self,
+ provider_id: str,
+ frame: common_pb2.DataFrame,
+ client_request_id: str | None = None,
+ ) -> None:
+ """
+ :param provider_id: The id registerProvider() returned. Required, non-blank, at most MAX_BUDGETED_ID_CHARS.
+ :param frame: The data to ingest.
+ :param client_request_id: Identifies this request in its ack and its request-status document. Defaults to
+ a generated uuid4 string: the server does not check uniqueness, and a reused id makes status lookups
+ ambiguous. Pass your own when it should mean something; it must be non-blank and at most
+ MAX_BUDGETED_ID_CHARS characters.
+ :raises ValueError: if an id is blank or too long, or the frame fails validation.
+ """
+ _require_id(provider_id, "provider_id")
+ if client_request_id is None:
+ client_request_id = str(uuid.uuid4())
+ else:
+ _require_id(client_request_id, "client_request_id")
+ validate_data_frame(frame, caller="IngestDataRequestParams")
+
+ self.provider_id = provider_id
+ self.frame = frame
+ self.client_request_id = client_request_id
+
+
+def chunked_request_params(
+ provider_id: str,
+ frames: Iterable[common_pb2.DataFrame],
+ base_request_id: str | None = None,
+) -> Iterator[IngestDataRequestParams]:
+ """
+ Wraps a sequence of frames -- typically split_data_frame()'s chunks -- as request params with correlated ids.
+
+ Request n (from 0) gets client_request_id "-", so the chunks of one frame are recognizable in their
+ acks and status documents. Lazy, like split_data_frame(), so it can feed ingest_data_stream() without holding
+ every chunk.
+
+ :param provider_id: The id registerProvider() returned.
+ :param frames: The frames to wrap.
+ :param base_request_id: The id prefix. Defaults to a generated uuid4 string. Must be non-blank, and short
+ enough to leave room for the "-" suffix within MAX_BUDGETED_ID_CHARS.
+ :return: An iterator over IngestDataRequestParams.
+ :raises ValueError: on the first iteration, if base_request_id is blank or too long; later, if a frame fails
+ validation or a suffixed id would exceed the cap.
+ """
+ base = base_request_id if base_request_id is not None else str(uuid.uuid4())
+ # Checked with the shortest suffix, before any frame is consumed, so a bad base fails as itself rather than as
+ # the first chunk's client_request_id. (A generator body runs at first next(), not at the call.)
+ _require_id(base, "base_request_id")
+ if len(base) + len("-0") > MAX_BUDGETED_ID_CHARS:
+ raise ValueError(
+ f"base_request_id is {len(base)} characters; with the '-' suffix the request ids would exceed "
+ f"{MAX_BUDGETED_ID_CHARS} characters"
+ )
+ for index, frame in enumerate(frames):
+ yield IngestDataRequestParams(provider_id, frame, client_request_id=f"{base}-{index}")
+
+
+class IngestDataApiResult(ApiResultBase):
+ """
+ Wraps one ingestion request's ack or reject.
+
+ An ack (is_error False) means only that the request passed validation; see query_request_status() for whether
+ it was ingested. A reject carries the server's message. provider_id and client_request_id identify the
+ request either way.
+ """
+
+ def __init__(
+ self,
+ is_error: bool,
+ message: str,
+ response: ingestion_pb2.IngestDataResponse | None = None,
+ ) -> None:
+ """
+ :param is_error: True for a reject or a failed call.
+ :param message: The error message, or "" on an ack.
+ :param response: The IngestDataResponse, or None if the call failed or the result was built by the unary
+ dispatch path for a reject (which does not keep the response).
+ """
+ super().__init__(is_error, message)
+ self.response = response
+ self.provider_id: str | None = response.providerId if response is not None else None
+ self.client_request_id: str | None = response.clientRequestId if response is not None else None
+
+ @property
+ def num_rows(self) -> int | None:
+ """Rows the server counted in the frame, echoed on an ack; None otherwise."""
+ if self.response is None or not self.response.HasField("ackResult"):
+ return None
+ return self.response.ackResult.numRows
+
+ @property
+ def num_columns(self) -> int | None:
+ """Columns the server counted in the frame, echoed on an ack; None otherwise."""
+ if self.response is None or not self.response.HasField("ackResult"):
+ return None
+ return self.response.ackResult.numColumns
+
+
+class IngestDataStreamApiResult(ApiResultBase):
+ """
+ Wraps ingestDataStream()'s single response.
+
+ **An error result keeps its response.** If any request in the stream was rejected, the server answers with an
+ ExceptionalResult ("one or more requests were rejected") -- but every request it did not reject was still
+ accepted, and rejected_request_ids is the only record of which failed. num_requests is None in that case,
+ since the server omits it. A failed call has no response, and all three properties are None.
+ """
+
+ def __init__(
+ self,
+ is_error: bool,
+ message: str,
+ response: ingestion_pb2.IngestDataStreamResponse | None = None,
+ ) -> None:
+ """
+ :param is_error: True if any request was rejected, or the call failed.
+ :param message: The error message, or "" on success.
+ :param response: The IngestDataStreamResponse, kept on error too; None only if the call failed.
+ """
+ super().__init__(is_error, message)
+ self.response = response
+
+ @property
+ def client_request_ids(self) -> list[str] | None:
+ """Every request id the server received, rejected ones included; None if the call failed."""
+ return list(self.response.clientRequestIds) if self.response is not None else None
+
+ @property
+ def rejected_request_ids(self) -> list[str] | None:
+ """The ids of the requests the server rejected; None if the call failed."""
+ return list(self.response.rejectedRequestIds) if self.response is not None else None
+
+ @property
+ def num_requests(self) -> int | None:
+ """How many requests were accepted, when none was rejected; None otherwise."""
+ if self.response is None or not self.response.HasField("ingestDataStreamResult"):
+ return None
+ return self.response.ingestDataStreamResult.numRequests
+
+
+class _RequestFeed:
+ """
+ The request iterator handed to grpcio for a streaming ingest, which remembers why it stopped.
+
+ grpcio consumes a request iterator on its own thread, and if the iterator raises, it cancels the call and
+ reports UNKNOWN "Exception iterating requests!" -- the original exception is neither chained nor re-raised.
+ The caller's iterable (a generator over split_data_frame(), say) is exactly where a real failure happens, so
+ this records the exception for the sender to re-raise in its place, with a count of the requests already handed
+ to gRPC. That count is an upper bound, not a delivery receipt: the cancellation races the sends, so some of
+ those requests may never reach the server. Any that did are ingested and are not undone.
+ """
+
+ def __init__(self, requests: Iterable[IngestDataRequestParams], build: Callable, logger: logging.Logger) -> None:
+ """
+ :param requests: The caller's request params, consumed lazily.
+ :param build: Turns one IngestDataRequestParams into an IngestDataRequest.
+ :param logger: The client's logger.
+ """
+ self._requests = requests
+ self._build = build
+ self._logger = logger
+ self.sent = 0
+ self.error: Exception | None = None
+
+ def __iter__(self) -> Iterator[ingestion_pb2.IngestDataRequest]:
+ try:
+ for params in self._requests:
+ request = self._build(params)
+ self.sent += 1
+ yield request
+ except Exception as e:
+ self.error = e
+ raise
+
+ def reraise(self, rpc_error: grpc.RpcError, op_name: str) -> None:
+ """
+ Re-raises the recorded exception, chained from the RpcError it caused, if there is one.
+
+ The exception is raised as itself, type and traceback intact, so the caller can catch what their own code
+ raised. The sent count travels as an exception note where the interpreter supports them (3.11+) and is
+ always logged.
+
+ :param rpc_error: The error grpcio reported.
+ :param op_name: The API method name, for the message.
+ :raises Exception: the recorded exception, if any.
+ """
+ if self.error is None:
+ return
+ note = (
+ f"{op_name}: raised while producing request {self.sent + 1}; {self.sent} request(s) had been handed to "
+ f"gRPC, and any that reached the server are ingested and not rolled back -- check request status"
+ )
+ self._logger.error("%s (%s: %s)", note, type(self.error).__name__, self.error)
+ if hasattr(self.error, "add_note"):
+ self.error.add_note(note)
+ raise self.error from rpc_error
+
+
+# ----------------------------------------------------------------------
+# request status
+# ----------------------------------------------------------------------
+
+
+class IngestionRequestStatus(IntEnum):
+ """
+ The outcome recorded in a request-status document, mirroring the proto enum so code compares names.
+
+ Beware the proto's zero value is SUCCESS: an unset or defaulted status field reads as success. Only a status
+ read from an actual document means anything.
+ """
+
+ SUCCESS = ingestion_pb2.INGESTION_REQUEST_STATUS_SUCCESS
+ REJECTED = ingestion_pb2.INGESTION_REQUEST_STATUS_REJECTED
+ ERROR = ingestion_pb2.INGESTION_REQUEST_STATUS_ERROR
+
+
+class RequestStatusQuery:
+ """
+ Factory of helpers building QueryRequestStatusRequest criteria for IngestionClient.query_request_status().
+ Criteria are ANDed. Each helper rejects the inputs the server would reject, naming the argument.
+
+ Note a RequestIdCriterion holds ONE id, and two of them AND to nothing: to check several requests, query by
+ provider and time range and match the ids yourself -- which is what await_request_statuses() does.
+
+ Example:
+ from dp_python_lib.client import RequestStatusQuery as RS, IngestionRequestStatus
+ criteria = [RS.provider_id(pid), RS.status([IngestionRequestStatus.ERROR]), RS.time_range(since)]
+ """
+
+ _Criterion = ingestion_pb2.QueryRequestStatusRequest.QueryRequestStatusCriterion
+
+ @staticmethod
+ def provider_id(provider_id: str) -> "ingestion_pb2.QueryRequestStatusRequest.QueryRequestStatusCriterion":
+ """
+ Matches documents for one provider, by the id registerProvider() returned.
+ :param provider_id: The provider id.
+ :return: A criterion with providerIdCriterion set.
+ :raises ValueError: if provider_id is blank.
+ """
+ if not provider_id or not provider_id.strip():
+ raise ValueError("provider_id() requires a non-blank provider_id")
+ criterion = RequestStatusQuery._Criterion()
+ criterion.providerIdCriterion.providerId = provider_id
+ return criterion
+
+ @staticmethod
+ def provider_name(provider_name: str) -> "ingestion_pb2.QueryRequestStatusRequest.QueryRequestStatusCriterion":
+ """
+ Matches documents for one provider, by the name it registered with.
+ :param provider_name: The provider name.
+ :return: A criterion with providerNameCriterion set.
+ :raises ValueError: if provider_name is blank.
+ """
+ if not provider_name or not provider_name.strip():
+ raise ValueError("provider_name() requires a non-blank provider_name")
+ criterion = RequestStatusQuery._Criterion()
+ criterion.providerNameCriterion.providerName = provider_name
+ return criterion
+
+ @staticmethod
+ def request_id(client_request_id: str) -> "ingestion_pb2.QueryRequestStatusRequest.QueryRequestStatusCriterion":
+ """
+ Matches documents for one clientRequestId. Ids are not unique server-side, so several may match.
+ :param client_request_id: The request's clientRequestId.
+ :return: A criterion with requestIdCriterion set.
+ :raises ValueError: if client_request_id is blank.
+ """
+ if not client_request_id or not client_request_id.strip():
+ raise ValueError("request_id() requires a non-blank client_request_id")
+ criterion = RequestStatusQuery._Criterion()
+ criterion.requestIdCriterion.requestId = client_request_id
+ return criterion
+
+ @staticmethod
+ def status(
+ statuses: Sequence[IngestionRequestStatus | int],
+ ) -> "ingestion_pb2.QueryRequestStatusRequest.QueryRequestStatusCriterion":
+ """
+ Matches documents whose status is any of the given ones.
+ :param statuses: IngestionRequestStatus members (or their int values).
+ :return: A criterion with statusCriterion set.
+ :raises ValueError: if statuses is empty or holds a value that is not a known status.
+ """
+ if not statuses:
+ raise ValueError("status() requires at least one status")
+ values: list[ingestion_pb2.IngestionRequestStatus.ValueType] = []
+ for status in statuses:
+ try:
+ # ValueType is int at runtime and a NewType under typed stubs, where a bare int would not check.
+ values.append(ingestion_pb2.IngestionRequestStatus.ValueType(int(IngestionRequestStatus(status))))
+ except ValueError:
+ raise ValueError(f"status() received an unknown status {status!r}") from None
+ criterion = RequestStatusQuery._Criterion()
+ criterion.statusCriterion.status[:] = values
+ return criterion
+
+ @staticmethod
+ def time_range(
+ begin: TimestampInput, end: TimestampInput | None = None
+ ) -> "ingestion_pb2.QueryRequestStatusRequest.QueryRequestStatusCriterion":
+ """
+ Matches documents created in [begin, end], both ends inclusive, at millisecond resolution.
+
+ The time is the status document's creation -- when the ingestion job FINISHED -- not when the request was
+ sent or the time of its data.
+ :param begin: Earliest creation time. Must be at least 1 s after the epoch.
+ :param end: Latest creation time; omit for "now" on the server's clock.
+ :return: A criterion with timeRangeCriterion set.
+ :raises ValueError: if begin is under 1 s after the epoch, or end is before begin.
+ """
+ begin_ts = to_timestamp(begin)
+ if to_epoch_nanos(begin_ts) < _MIN_STATUS_QUERY_BEGIN_NANOS:
+ raise ValueError("time_range() requires begin at least 1 second after the epoch")
+ criterion = RequestStatusQuery._Criterion()
+ criterion.timeRangeCriterion.beginTime.CopyFrom(begin_ts)
+ if end is not None:
+ end_ts = to_timestamp(end)
+ if to_epoch_nanos(end_ts) < to_epoch_nanos(begin_ts):
+ raise ValueError("time_range() requires end at or after begin")
+ criterion.timeRangeCriterion.endTime.CopyFrom(end_ts)
+ return criterion
+
+
+class QueryRequestStatusApiResult(ApiResultBase):
+ """
+ Wraps queryRequestStatus(): the matching request-status documents, in document-id (roughly creation) order.
+ """
+
+ def __init__(
+ self,
+ is_error: bool,
+ message: str,
+ response: ingestion_pb2.QueryRequestStatusResponse | None = None,
+ ) -> None:
+ """
+ :param is_error: True if the query was rejected or the call failed.
+ :param message: The error message, or "" on success.
+ :param response: The QueryRequestStatusResponse, or None on error.
+ """
+ super().__init__(is_error, message)
+ self.response = response
+
+ @property
+ def request_statuses(
+ self,
+ ) -> "list[ingestion_pb2.QueryRequestStatusResponse.RequestStatusResult.RequestStatus] | None":
+ """The matching documents; None on error."""
+ if self.response is None or not self.response.HasField("requestStatusResult"):
+ return None
+ return list(self.response.requestStatusResult.requestStatus)
+
+
+# ----------------------------------------------------------------------
+# client
+# ----------------------------------------------------------------------
+
class IngestionClient(ServiceApiClientBase):
"""
This is the user-facing Ingestion Service API class. It provides methods and utility classes for calling
Ingestion Service methods.
+
+ The typical sequence: register_provider() once for the provider id; ingest with ingest_data(),
+ ingest_data_stream(), or iter_ingest_data_bidi_stream(); then confirm with await_request_statuses(), because
+ an ack means only that a request passed validation.
"""
def __init__(self, channel: grpc.Channel) -> None:
@@ -65,6 +593,8 @@ def __init__(self, channel: grpc.Channel) -> None:
self.logger = logging.getLogger(__name__)
self.logger.debug("IngestionClient initialized with channel: %s", channel)
+ # ---- registerProvider
+
def _build_register_provider_request(
self, request_params: RegisterProviderRequestParams
) -> ingestion_pb2.RegisterProviderRequest:
@@ -124,7 +654,8 @@ def register_provider(self, request_params: RegisterProviderRequestParams) -> Re
"""
User facing method for invoking the registerProvider() API method.
:param request_params: Contains user parameters for call to registerProvider() API method.
- :return: Returns RegisterProviderApiResult with the method response and status information.
+ :return: Returns RegisterProviderApiResult with the method response and status information; its
+ provider_id is what every ingestion request needs.
"""
self.logger.info("Starting registerProvider operation for provider: %s", request_params.name)
@@ -140,3 +671,358 @@ def register_provider(self, request_params: RegisterProviderRequestParams) -> Re
)
return result
+
+ # ---- ingestData
+
+ def _build_ingest_data_request(self, request_params: IngestDataRequestParams) -> ingestion_pb2.IngestDataRequest:
+ """
+ Builds an IngestDataRequest from validated params.
+ :param request_params: The request's params.
+ :return: The IngestDataRequest.
+ """
+ request = ingestion_pb2.IngestDataRequest()
+ request.providerId = request_params.provider_id
+ request.clientRequestId = request_params.client_request_id
+ request.ingestionDataFrame.CopyFrom(request_params.frame)
+ return request
+
+ def _send_ingest_data(self, request: ingestion_pb2.IngestDataRequest) -> IngestDataApiResult:
+ """
+ Invokes the unary ingestData() API method.
+ :param request: The IngestDataRequest.
+ :return: An IngestDataApiResult; the ids are filled in by ingest_data() when the response is not kept.
+ """
+ return self._dispatch(
+ self._stub.ingestData,
+ request,
+ IngestDataApiResult,
+ "ackResult",
+ "ingestData",
+ request_log=lambda: self.logger.info(
+ "Calling ingestData API for request %s (provider %s)", request.clientRequestId, request.providerId
+ ),
+ success_log=lambda response: self.logger.info(
+ "ingestData acked request %s: %d rows x %d columns",
+ response.clientRequestId,
+ response.ackResult.numRows,
+ response.ackResult.numColumns,
+ ),
+ rpc_error_hint=_ingest_size_hint,
+ )
+
+ def ingest_data(self, request_params: IngestDataRequestParams) -> IngestDataApiResult:
+ """
+ Sends one ingestion request and returns its ack or reject.
+
+ An ack means the request passed validation and was queued -- NOT that it was ingested. Use
+ await_request_statuses() (or query_request_status()) to learn that.
+
+ A request over the server's inbound message limit (4,096,000 bytes by default) fails as a gRPC error; the
+ message then says so and points at split_data_frame().
+
+ :param request_params: The request (see IngestDataRequestParams).
+ :return: An IngestDataApiResult whose provider_id and client_request_id identify the request.
+ """
+ request = self._build_ingest_data_request(request_params)
+ result = self._send_ingest_data(request)
+ # _dispatch keeps no response on an error, so name the request from what was sent.
+ result.provider_id = request_params.provider_id
+ result.client_request_id = request_params.client_request_id
+ if result.result_status.is_error:
+ self.logger.error(
+ "ingestData request %s failed: %s", request_params.client_request_id, result.result_status.message
+ )
+ return result
+
+ # ---- ingestDataStream
+
+ def _send_ingest_data_stream(self, feed: _RequestFeed) -> IngestDataStreamApiResult:
+ """
+ Invokes the client-streaming ingestDataStream() API method.
+
+ Hand-written rather than _dispatch, because an error result must KEEP its response: rejectedRequestIds
+ arrives on the ExceptionalResult response and is the only record of which requests failed. And the feed's
+ own exception, if any, is re-raised in place of grpcio's generic error.
+
+ :param feed: The request iterator.
+ :return: An IngestDataStreamApiResult.
+ :raises Exception: whatever the caller's request iterable raised.
+ """
+ self.logger.info("Calling ingestDataStream API")
+ try:
+ response = self._stub.ingestDataStream(iter(feed))
+ except grpc.RpcError as e:
+ feed.reraise(e, "ingestDataStream")
+ error_msg = f"gRPC error: {e.details()}{_ingest_size_hint(e)}"
+ self.logger.error("gRPC error during ingestDataStream after %d request(s): %s", feed.sent, e.details())
+ return IngestDataStreamApiResult(is_error=True, message=error_msg)
+ except Exception as e:
+ error_msg = f"Unexpected error: {e!s}"
+ self.logger.exception("Unexpected error during ingestDataStream: %s", str(e))
+ return IngestDataStreamApiResult(is_error=True, message=error_msg)
+
+ if response.HasField("exceptionalResult"):
+ error_msg = response.exceptionalResult.message
+ self.logger.warning(
+ "ingestDataStream rejected %d of %d request(s): %s",
+ len(response.rejectedRequestIds),
+ len(response.clientRequestIds),
+ error_msg,
+ )
+ return IngestDataStreamApiResult(is_error=True, message=error_msg, response=response)
+ if response.HasField("ingestDataStreamResult"):
+ self.logger.info("ingestDataStream accepted %d request(s)", response.ingestDataStreamResult.numRequests)
+ return IngestDataStreamApiResult(is_error=False, message="", response=response)
+ error_msg = "Unexpected response format: neither exceptionalResult nor ingestDataStreamResult found"
+ self.logger.error(error_msg)
+ return IngestDataStreamApiResult(is_error=True, message=error_msg, response=response)
+
+ def ingest_data_stream(self, requests: Iterable[IngestDataRequestParams]) -> IngestDataStreamApiResult:
+ """
+ Sends many ingestion requests on one client-streaming call and returns the server's single summary.
+
+ The iterable is consumed lazily, so a generator over split_data_frame() (see chunked_request_params())
+ never holds every request at once. The server validates each request as it arrives; a reject does not end
+ the stream, and every valid request is queued for ingestion.
+
+ - All accepted: is_error False, num_requests set.
+ - Some rejected: is_error True, rejected_request_ids naming them. This does NOT raise, because the other
+ requests were accepted and raising would hide which.
+ - If the iterable itself raises, that exception is re-raised as itself (grpcio would otherwise report only
+ "Exception iterating requests!"). Requests sent before it may or may not have reached the server; any
+ that did stay ingested, so check their status.
+
+ As with ingest_data(), acceptance is not ingestion; confirm with await_request_statuses().
+
+ :param requests: The requests to send.
+ :return: An IngestDataStreamApiResult.
+ :raises Exception: whatever the iterable raised.
+ """
+ self.logger.info("Starting ingestDataStream operation")
+ feed = _RequestFeed(requests, self._build_ingest_data_request, self.logger)
+ return self._send_ingest_data_stream(feed)
+
+ # ---- ingestDataBidiStream
+
+ def _send_ingest_data_bidi_stream(self, feed: _RequestFeed) -> Generator[IngestDataApiResult, None, None]:
+ """
+ Invokes the bidirectional ingestDataBidiStream() API method, yielding one result per response.
+
+ A per-request reject is yielded with its response, which names the request. A transport or unexpected
+ error is yielded WITHOUT a response and ends the generator; that absence is how the public wrapper tells
+ the two apart. The feed's own exception, if any, is re-raised instead.
+
+ :param feed: The request iterator.
+ :return: An iterator over results.
+ :raises Exception: whatever the caller's request iterable raised.
+ """
+ self.logger.info("Calling ingestDataBidiStream API")
+ call = None
+ try:
+ call = self._stub.ingestDataBidiStream(iter(feed))
+ for response in call:
+ if response.HasField("exceptionalResult"):
+ error_msg = response.exceptionalResult.message
+ self.logger.warning(
+ "ingestDataBidiStream rejected request %s: %s", response.clientRequestId, error_msg
+ )
+ yield IngestDataApiResult(is_error=True, message=error_msg, response=response)
+ elif response.HasField("ackResult"):
+ yield IngestDataApiResult(is_error=False, message="", response=response)
+ else:
+ error_msg = "Unexpected response format: neither exceptionalResult nor ackResult found"
+ self.logger.error("%s (request %s)", error_msg, response.clientRequestId)
+ yield IngestDataApiResult(is_error=True, message=error_msg, response=response)
+ except grpc.RpcError as e:
+ feed.reraise(e, "ingestDataBidiStream")
+ error_msg = f"gRPC error: {e.details()}{_ingest_size_hint(e)}"
+ self.logger.error("gRPC error during ingestDataBidiStream after %d request(s): %s", feed.sent, e.details())
+ yield IngestDataApiResult(is_error=True, message=error_msg)
+ except Exception as e:
+ error_msg = f"Unexpected error: {e!s}"
+ self.logger.exception("Unexpected error during ingestDataBidiStream: %s", str(e))
+ yield IngestDataApiResult(is_error=True, message=error_msg)
+ finally:
+ # A caller that stops reading early (break, or an exception in its loop) closes this generator, but
+ # grpcio would keep pulling and SENDING requests on its own thread -- ingesting data the caller believes
+ # it abandoned. grpcio's call object does cancel itself when garbage-collected, but that ties the
+ # behavior to collection timing; cancel explicitly. On a call that already finished it does nothing.
+ if call is not None and hasattr(call, "cancel"):
+ call.cancel()
+
+ def iter_ingest_data_bidi_stream(
+ self, requests: Iterable[IngestDataRequestParams]
+ ) -> Iterator[IngestDataApiResult]:
+ """
+ Sends ingestion requests on a bidirectional stream, yielding each one's ack or reject as it arrives.
+
+ One result per request, in order. **A reject is yielded as an is_error result, not raised**: it is a fact
+ about one request, and the server keeps processing the rest. This deliberately differs from
+ iter_query_samples_stream(), where an error ends the stream. A transport error does end it, as a
+ RuntimeError. If the request iterable itself raises, that exception is re-raised as itself; as with
+ ingest_data_stream(), requests sent before it may have been ingested.
+
+ grpcio consumes the iterable on its own thread, so a producer generator runs concurrently with the
+ responses being read. The server applies backpressure through HTTP/2 flow control.
+
+ **Closing the returned iterator cancels the call**, so no further requests are sent (requests already sent
+ may still be ingested). A bare `break` does NOT close it while anything still references it: until it is
+ closed or garbage-collected, grpcio keeps pulling and sending requests. So to stop early, close it
+ explicitly -- `contextlib.closing` does that on every exit, a break or an exception included:
+
+ with contextlib.closing(ingestion.iter_ingest_data_bidi_stream(requests)) as results:
+ for result in results:
+ if result.result_status.is_error:
+ break # the call is cancelled as the with block exits
+
+ Reading to the end needs no close: the call has finished by then.
+
+ :param requests: The requests to send.
+ :return: A lazy iterator over per-request results.
+ :raises RuntimeError: on a transport or unexpected error.
+ :raises Exception: whatever the iterable raised.
+ """
+ self.logger.info("Starting ingestDataBidiStream operation")
+ feed = _RequestFeed(requests, self._build_ingest_data_request, self.logger)
+ # closing() so that closing THIS generator closes the sender's at once, running the finally that cancels the
+ # call. It does not make a caller's `break` a close: a suspended generator that is still referenced is
+ # closed only explicitly or when collected, which is why the docstring asks for contextlib.closing.
+ with contextlib.closing(self._send_ingest_data_bidi_stream(feed)) as results:
+ for result in results:
+ if result.result_status.is_error and result.response is None:
+ raise RuntimeError(f"ingestDataBidiStream failed: {result.result_status.message}")
+ yield result
+
+ # ---- queryRequestStatus
+
+ def _send_query_request_status(
+ self, request: ingestion_pb2.QueryRequestStatusRequest
+ ) -> QueryRequestStatusApiResult:
+ """
+ Invokes the queryRequestStatus() API method.
+ :param request: The QueryRequestStatusRequest.
+ :return: A QueryRequestStatusApiResult.
+ """
+ return self._dispatch(
+ self._stub.queryRequestStatus,
+ request,
+ QueryRequestStatusApiResult,
+ "requestStatusResult",
+ "queryRequestStatus",
+ request_log=lambda: self.logger.info(
+ "Calling queryRequestStatus API with %d criteria", len(request.criteria)
+ ),
+ success_log=lambda response: self.logger.info(
+ "queryRequestStatus returned %d documents", len(response.requestStatusResult.requestStatus)
+ ),
+ rpc_error_hint=_status_query_size_hint,
+ )
+
+ def query_request_status(
+ self,
+ criteria: "Sequence[ingestion_pb2.QueryRequestStatusRequest.QueryRequestStatusCriterion]",
+ ) -> QueryRequestStatusApiResult:
+ """
+ Queries request-status documents: the only record of whether an acked request was actually ingested.
+
+ Each handled request gets one document, written after its data (so SUCCESS means persisted and queryable),
+ or recording REJECTED / ERROR. A few paths write none -- a reject while the service is shutting down, and
+ a handler that fails outright -- so absence is "unknown", never success.
+
+ The server returns every match in one message, with no paging yet, so keep queries narrow: a provider plus
+ a time range, not a bare status.
+
+ :param criteria: At least one criterion (see RequestStatusQuery); they are ANDed.
+ :return: A QueryRequestStatusApiResult.
+ :raises ValueError: if criteria is empty (the server rejects it).
+ """
+ if not criteria:
+ raise ValueError("query_request_status() requires at least one criterion")
+ request = ingestion_pb2.QueryRequestStatusRequest()
+ request.criteria.extend(criteria)
+ return self._send_query_request_status(request)
+
+ def await_request_statuses(
+ self,
+ provider_id: str,
+ client_request_ids: Iterable[str],
+ *,
+ since: TimestampInput,
+ timeout: float = 30.0,
+ poll_interval: float = 0.25,
+ ) -> "dict[str, list[ingestion_pb2.QueryRequestStatusResponse.RequestStatusResult.RequestStatus]]":
+ """
+ Polls until every given request has a status document, and returns them.
+
+ This does not judge success: check each document's ingestionRequestStatus against IngestionRequestStatus.
+ A request can be acked and still end in ERROR (an unknown providerId, a duplicate PV + first timestamp).
+
+ Each poll is ONE query -- this provider, documents created since `since` -- with the ids matched here, since
+ the API cannot ask about several request ids at once. `since` is required, and should be captured before
+ the first request is sent:
+
+ since = datetime.now(timezone.utc)
+ result = ingestion.ingest_data(params)
+ statuses = ingestion.await_request_statuses(provider_id, [result.client_request_id], since=since)
+
+ It bounds the query (the method has no paging, and a provider's whole history can exceed the receive limit),
+ and it keeps out documents from earlier runs that reused an id. The floor is backed off by
+ REQUEST_STATUS_CLOCK_SKEW for the server's clock, so a reused id can still match a document from within that
+ margin; the default generated ids avoid that.
+
+ :param provider_id: The provider the requests were sent as.
+ :param client_request_ids: The requests to wait for.
+ :param since: A time at or before the first request was sent.
+ :param timeout: Seconds to wait before giving up.
+ :param poll_interval: Seconds between polls.
+ :return: {client_request_id: [documents]}, one entry per distinct id; a list because ids are not unique.
+ :raises ValueError: on blank or empty arguments, or a non-positive timeout or poll_interval.
+ :raises TimeoutError: if some ids have no document by the timeout; the message names them.
+ :raises RuntimeError: if a status query fails.
+ """
+ if not provider_id or not provider_id.strip():
+ raise ValueError("await_request_statuses() requires a non-blank provider_id")
+ wanted = list(dict.fromkeys(client_request_ids))
+ if not wanted:
+ raise ValueError("await_request_statuses() requires at least one client_request_id")
+ for request_id in wanted:
+ if not request_id or not request_id.strip():
+ raise ValueError(f"await_request_statuses() received a blank client_request_id: {request_id!r}")
+ if timeout <= 0 or poll_interval <= 0:
+ raise ValueError("await_request_statuses() requires a positive timeout and poll_interval")
+
+ # Floor division of two timedeltas is exact integer arithmetic.
+ skew_nanos = (REQUEST_STATUS_CLOCK_SKEW // timedelta(microseconds=1)) * 1_000
+ floor_nanos = max(to_epoch_nanos(to_timestamp(since)) - skew_nanos, _MIN_STATUS_QUERY_BEGIN_NANOS)
+ criteria = [
+ RequestStatusQuery.provider_id(provider_id),
+ RequestStatusQuery.time_range(from_epoch_nanos(floor_nanos)),
+ ]
+
+ wanted_set = set(wanted)
+ deadline = time.monotonic() + timeout
+ attempt = 0
+ while True:
+ attempt += 1
+ result = self.query_request_status(criteria)
+ if result.result_status.is_error:
+ raise RuntimeError(f"await_request_statuses() status query failed: {result.result_status.message}")
+ found: dict[str, list] = {}
+ for document in result.request_statuses or []:
+ if document.requestId in wanted_set:
+ found.setdefault(document.requestId, []).append(document)
+ missing = [request_id for request_id in wanted if request_id not in found]
+ if not missing:
+ self.logger.info(
+ "await_request_statuses: all %d request(s) have a status (poll %d)", len(wanted), attempt
+ )
+ return {request_id: found[request_id] for request_id in wanted}
+
+ remaining = deadline - time.monotonic()
+ if remaining <= 0:
+ shown = ", ".join(missing[:10]) + (f", ... ({len(missing) - 10} more)" if len(missing) > 10 else "")
+ raise TimeoutError(
+ f"await_request_statuses() timed out after {timeout} s with {len(missing)} of {len(wanted)} "
+ f"request(s) still lacking a status document: {shown}"
+ )
+ time.sleep(min(poll_interval, remaining))
diff --git a/src/dp_python_lib/client/sample_status_conversions.py b/src/dp_python_lib/client/sample_status_conversions.py
index dee8fef..838d07a 100644
--- a/src/dp_python_lib/client/sample_status_conversions.py
+++ b/src/dp_python_lib/client/sample_status_conversions.py
@@ -21,7 +21,7 @@
from collections.abc import Iterator
-from dp_python_lib.client.time_conversions import NANOS_PER_SECOND, to_epoch_nanos
+from dp_python_lib.client.time_conversions import from_epoch_nanos, to_epoch_nanos
from dp_python_lib.grpc import common_pb2
# The Timestamp -> epoch-nanoseconds conversion is shared (see time_conversions.to_epoch_nanos); this private
@@ -29,15 +29,8 @@
_timestamp_to_nanos = to_epoch_nanos
-def _nanos_to_timestamp(epoch_nanos: int) -> common_pb2.Timestamp:
- """
- Converts integer epoch nanoseconds back into a common.Timestamp.
- :param epoch_nanos: Epoch nanoseconds.
- :return: The equivalent common.Timestamp.
- """
- timestamp = common_pb2.Timestamp()
- timestamp.epochSeconds, timestamp.nanoseconds = divmod(epoch_nanos, NANOS_PER_SECOND)
- return timestamp
+# Likewise the inverse (time_conversions.from_epoch_nanos).
+_nanos_to_timestamp = from_epoch_nanos
def expand_data_timestamps(data_timestamps: common_pb2.DataTimestamps) -> list[int]:
diff --git a/src/dp_python_lib/client/service_api_client_base.py b/src/dp_python_lib/client/service_api_client_base.py
index 5f2775c..0b8911e 100644
--- a/src/dp_python_lib/client/service_api_client_base.py
+++ b/src/dp_python_lib/client/service_api_client_base.py
@@ -50,6 +50,7 @@ def _dispatch(
op_name: str,
request_log: Callable[[], None] | None = None,
success_log: Callable[[Any], None] | None = None,
+ rpc_error_hint: Callable[[grpc.RpcError], str] | None = None,
) -> ApiResultT:
"""
Invokes a unary gRPC API method and applies the standard three-tier error handling shared by every unary
@@ -76,6 +77,10 @@ def _dispatch(
generic "Calling API" is logged instead.
:param success_log: Optional callable, passed the response, logging a method-specific success message (some
methods report a record count read off the response). When omitted, a generic message is logged instead.
+ :param rpc_error_hint: Optional callable, passed a caught grpc.RpcError, returning text to APPEND to the
+ "gRPC error: " message -- advice that grpcio's bare details cannot give, such as pointing an
+ oversized ingest at split_data_frame(). Return "" to add nothing. It is appended, never substituted,
+ because the message prefix is part of the result contract.
:return: A result_cls instance with the method response and status information.
"""
if request_log is not None:
@@ -111,6 +116,8 @@ def _dispatch(
except grpc.RpcError as e:
error_msg = f"gRPC error: {e.details()}"
+ if rpc_error_hint is not None:
+ error_msg += rpc_error_hint(e)
# Safely get the error code -- it may not be available on a bare RpcError, as raised by test mocks.
try:
error_code = e.code()
diff --git a/src/dp_python_lib/client/time_conversions.py b/src/dp_python_lib/client/time_conversions.py
index 32cfb64..11e51b9 100644
--- a/src/dp_python_lib/client/time_conversions.py
+++ b/src/dp_python_lib/client/time_conversions.py
@@ -3,7 +3,8 @@
Every API that takes an instant accepts the same three spellings -- a timezone-aware datetime, epoch seconds, or an
already-built common.Timestamp -- and every conversions module that reads one back wants integer nanoseconds. Those
-two directions are `to_timestamp()` and `to_epoch_nanos()`, and they live here rather than in a feature client.
+two directions are `to_timestamp()` and `to_epoch_nanos()`, and they live here rather than in a feature client, along
+with `from_epoch_nanos()`, which turns integer nanoseconds back into a Timestamp.
They were originally defined in machine_config_client, the first module to need them, and four other modules grew
imports of `to_timestamp()` from there -- which read as though datasets, queries, and DataFrames depended on the
@@ -52,6 +53,22 @@ def to_epoch_nanos(timestamp: common_pb2.Timestamp) -> int:
return timestamp.epochSeconds * NANOS_PER_SECOND + timestamp.nanoseconds
+def from_epoch_nanos(epoch_nanos: int) -> common_pb2.Timestamp:
+ """
+ Converts integer epoch nanoseconds into a common.Timestamp -- the inverse of to_epoch_nanos().
+
+ Integer division throughout, for the same exactness reason as to_epoch_nanos(). It was written out privately
+ in data_frame_conversions and sample_status_conversions before split_data_frame() became a third caller.
+
+ :param epoch_nanos: Epoch nanoseconds, as an int.
+ :return: The equivalent common.Timestamp.
+ :raises ValueError: if epoch_nanos is negative, since Timestamp.epochSeconds is unsigned.
+ """
+ timestamp = common_pb2.Timestamp()
+ timestamp.epochSeconds, timestamp.nanoseconds = divmod(epoch_nanos, NANOS_PER_SECOND)
+ return timestamp
+
+
def to_timestamp(value: TimestampInput) -> common_pb2.Timestamp:
"""
Converts a user-supplied time value into a common.Timestamp{epochSeconds, nanoseconds}.
diff --git a/tests/unit/test_data_frame.py b/tests/unit/test_data_frame.py
index e888475..386de94 100644
--- a/tests/unit/test_data_frame.py
+++ b/tests/unit/test_data_frame.py
@@ -7,7 +7,7 @@
sys.path.insert(0, os.path.join(os.path.dirname(__file__), "../../src"))
from dp_python_lib.client import data_frame as dfb
-from dp_python_lib.grpc import common_pb2
+from dp_python_lib.grpc import common_pb2, ingestion_pb2
try:
import numpy as np
@@ -25,6 +25,18 @@ def _axis(count=3):
return dfb.sampling_clock(T0, 1_000_000_000, count)
+def _hand_built_image_column(images, name="camera"):
+ """An ImageColumn built from the proto directly, with a valid descriptor."""
+ column = common_pb2.ImageColumn()
+ column.name = name
+ column.imageDescriptor.width = 2
+ column.imageDescriptor.height = 2
+ column.imageDescriptor.channels = 1
+ column.imageDescriptor.encoding = "raw-u8"
+ column.images.extend(images)
+ return column
+
+
class TestSamplingClock(unittest.TestCase):
"""The axis builders relocated here from sample_status_client in Phase 2; behavior is unchanged."""
@@ -384,7 +396,7 @@ def test_rejects_empty_axis(self):
dfb.data_frame(common_pb2.DataTimestamps(), [dfb.double_column("d", [1.0])])
def test_accepts_prebuilt_array_column(self):
- # Array builders are #17's; a hand-built one passes through, with dims giving the sample count.
+ # A hand-built array column passes through alongside the builders' output, with dims giving the count.
column = common_pb2.DoubleArrayColumn()
column.name = "waveform"
column.dimensions.dims.extend([4])
@@ -404,16 +416,13 @@ def test_array_column_count_uses_dims_product(self):
def test_accepts_prebuilt_image_column(self):
# ImageColumn keeps one payload per sample in `images`; it has no `values` field at all, so a generic
# len(column.values) count raises AttributeError on a column type data_frame() claims to support.
- column = common_pb2.ImageColumn()
- column.name = "camera"
- column.images.extend([b"frame-0", b"frame-1"])
+ column = _hand_built_image_column([b"frame-0", b"frame-1"])
frame = dfb.data_frame(_axis(2), [column])
self.assertEqual([c.name for c in frame.imageColumns], ["camera"])
def test_image_column_count_is_validated_against_the_axis(self):
- column = common_pb2.ImageColumn()
- column.name = "camera"
- column.images.extend([b"only-one"])
+ # A valid descriptor, so the count is what fails -- not the descriptor check that runs before it.
+ column = _hand_built_image_column([b"only-one"])
with self.assertRaises(ValueError) as ctx:
dfb.data_frame(_axis(3), [column])
self.assertIn("1 values", str(ctx.exception))
@@ -468,10 +477,513 @@ def test_serialized_column_gets_name_checks_only(self):
self.assertEqual([c.name for c in frame.serializedDataColumns], ["serialized"])
def test_serialized_column_still_needs_a_unique_name(self):
+ # Given an encoding, so the duplicate name is the only thing wrong with it.
column = common_pb2.SerializedDataColumn()
column.name = "dup"
- with self.assertRaises(ValueError):
+ column.encoding = "arrow"
+ with self.assertRaises(ValueError) as ctx:
dfb.data_frame(_axis(1), [dfb.double_column("dup", [1.0]), column])
+ self.assertIn("unique column names", str(ctx.exception))
+
+
+class TestStructuralColumnChecks(unittest.TestCase):
+ """
+ Each column kind's structural fields, checked on HAND-BUILT columns (#17, T7).
+
+ The builders enforce these rules too, but a column built from the proto bypasses every builder -- the recurring
+ #6 defect shape -- so the checks that matter are the ones data_frame() applies to any column.
+ """
+
+ def test_rejects_enum_column_without_enum_id(self):
+ for enum_id in ("", " "):
+ with self.subTest(enum_id=enum_id):
+ column = common_pb2.EnumColumn()
+ column.name = "state"
+ column.enumId = enum_id
+ column.values[:] = [0, 1]
+ with self.assertRaises(ValueError) as ctx:
+ dfb.data_frame(_axis(2), [column])
+ self.assertIn("enumId", str(ctx.exception))
+
+ def test_enum_builder_rejects_blank_enum_id(self):
+ with self.assertRaises(ValueError) as ctx:
+ dfb.enum_column("state", [0], enum_id=" ")
+ self.assertIn("non-blank enum_id", str(ctx.exception))
+
+ def test_rejects_struct_column_without_schema_id(self):
+ column = common_pb2.StructColumn()
+ column.name = "bpm"
+ column.values.extend([b"a", b"b"])
+ with self.assertRaises(ValueError) as ctx:
+ dfb.data_frame(_axis(2), [column])
+ self.assertIn("schemaId", str(ctx.exception))
+
+ def test_rejects_serialized_column_without_encoding(self):
+ for encoding in ("", "\t"):
+ with self.subTest(encoding=encoding):
+ column = common_pb2.SerializedDataColumn()
+ column.name = "blob"
+ column.encoding = encoding
+ with self.assertRaises(ValueError) as ctx:
+ dfb.data_frame(_axis(1), [dfb.double_column("d", [1.0]), column])
+ self.assertIn("encoding", str(ctx.exception))
+
+ def test_rejects_image_column_without_descriptor(self):
+ column = common_pb2.ImageColumn()
+ column.name = "camera"
+ column.images.extend([b"x"])
+ with self.assertRaises(ValueError) as ctx:
+ dfb.data_frame(_axis(1), [column])
+ self.assertIn("imageDescriptor", str(ctx.exception))
+
+ def test_rejects_image_descriptor_with_a_zero_dimension(self):
+ for field in ("width", "height", "channels"):
+ with self.subTest(field=field):
+ column = _hand_built_image_column([b"x"])
+ setattr(column.imageDescriptor, field, 0)
+ with self.assertRaises(ValueError) as ctx:
+ dfb.data_frame(_axis(1), [column])
+ self.assertIn(f"imageDescriptor.{field} > 0", str(ctx.exception))
+
+ def test_rejects_image_descriptor_with_blank_encoding(self):
+ column = _hand_built_image_column([b"x"])
+ column.imageDescriptor.encoding = " "
+ with self.assertRaises(ValueError) as ctx:
+ dfb.data_frame(_axis(1), [column])
+ self.assertIn("imageDescriptor.encoding", str(ctx.exception))
+
+ def test_rejects_array_column_with_more_than_three_dims(self):
+ column = common_pb2.DoubleArrayColumn()
+ column.name = "tensor"
+ column.dimensions.dims.extend([1, 1, 1, 2])
+ column.values[:] = [1.0, 2.0]
+ with self.assertRaises(ValueError) as ctx:
+ dfb.data_frame(_axis(1), [column])
+ self.assertIn("1 to 3", str(ctx.exception))
+
+ def test_accepts_three_dims(self):
+ column = common_pb2.DoubleArrayColumn()
+ column.name = "tensor"
+ column.dimensions.dims.extend([1, 1, 2])
+ column.values[:] = [1.0, 2.0]
+ frame = dfb.data_frame(_axis(1), [column])
+ self.assertEqual(len(frame.doubleArrayColumns), 1)
+
+
+class TestValidateDataFrame(unittest.TestCase):
+ """validate_data_frame() applies data_frame()'s checks to a frame assembled any other way."""
+
+ def test_accepts_a_frame_from_data_frame(self):
+ frame = dfb.data_frame(_axis(2), [dfb.double_column("d", [1.0, 2.0]), dfb.enum_column("e", [0, 1], "E")])
+ dfb.validate_data_frame(frame)
+
+ def test_rejects_a_hand_built_frame_with_a_hand_built_enum_lacking_its_id(self):
+ # The T7 point at frame level: nothing about a hand-built frame goes through data_frame().
+ frame = common_pb2.DataFrame()
+ frame.dataTimestamps.CopyFrom(_axis(2))
+ column = frame.enumColumns.add()
+ column.name = "state"
+ column.values[:] = [0, 1]
+ with self.assertRaises(ValueError) as ctx:
+ dfb.validate_data_frame(frame)
+ self.assertIn("enumId", str(ctx.exception))
+
+ def test_rejects_a_frame_with_no_columns(self):
+ frame = common_pb2.DataFrame()
+ frame.dataTimestamps.CopyFrom(_axis(2))
+ with self.assertRaises(ValueError) as ctx:
+ dfb.validate_data_frame(frame)
+ self.assertIn("at least one column", str(ctx.exception))
+
+ def test_rejects_a_frame_with_no_axis(self):
+ frame = common_pb2.DataFrame()
+ frame.doubleColumns.add(name="d").values[:] = [1.0]
+ with self.assertRaises(ValueError):
+ dfb.validate_data_frame(frame)
+
+ def test_rejects_duplicate_names_across_column_types(self):
+ frame = common_pb2.DataFrame()
+ frame.dataTimestamps.CopyFrom(_axis(1))
+ frame.doubleColumns.add(name="x").values[:] = [1.0]
+ frame.int32Columns.add(name="x").values[:] = [1]
+ with self.assertRaises(ValueError) as ctx:
+ dfb.validate_data_frame(frame)
+ self.assertIn("unique column names", str(ctx.exception))
+
+ def test_rejects_a_count_mismatch(self):
+ frame = common_pb2.DataFrame()
+ frame.dataTimestamps.CopyFrom(_axis(3))
+ frame.doubleColumns.add(name="d").values[:] = [1.0]
+ with self.assertRaises(ValueError) as ctx:
+ dfb.validate_data_frame(frame)
+ self.assertIn("1 values", str(ctx.exception))
+
+ def test_messages_name_the_function_the_caller_called(self):
+ # A frame that never went through data_frame() should not be reported as if it had.
+ frame = common_pb2.DataFrame()
+ frame.dataTimestamps.CopyFrom(_axis(2))
+ frame.enumColumns.add(name="state").values[:] = [0, 1]
+ for kwargs, prefix in (
+ ({}, "validate_data_frame() "),
+ ({"caller": "split_data_frame()"}, "split_data_frame() "),
+ ):
+ with self.subTest(prefix=prefix), self.assertRaises(ValueError) as ctx:
+ dfb.validate_data_frame(frame, **kwargs)
+ self.assertTrue(str(ctx.exception).startswith(prefix), str(ctx.exception))
+
+
+class TestArrayColumnBuilders(unittest.TestCase):
+ def test_one_dimensional_samples_infer_their_dims(self):
+ column = dfb.double_array_column("wf", [[1.0, 2.0, 3.0], [4.0, 5.0, 6.0]])
+ self.assertEqual(list(column.dimensions.dims), [3])
+ self.assertEqual(list(column.values), [1.0, 2.0, 3.0, 4.0, 5.0, 6.0])
+
+ def test_multidimensional_samples_flatten_row_major(self):
+ column = dfb.int32_array_column("m", [[[1, 2, 3], [4, 5, 6]]])
+ self.assertEqual(list(column.dimensions.dims), [2, 3])
+ self.assertEqual(list(column.values), [1, 2, 3, 4, 5, 6])
+
+ def test_three_dimensions_are_accepted_and_four_are_not(self):
+ column = dfb.int64_array_column("t", [[[[1, 2]]]])
+ self.assertEqual(list(column.dimensions.dims), [1, 1, 2])
+ with self.assertRaises(ValueError) as ctx:
+ dfb.int64_array_column("t", [[[[[1, 2]]]]])
+ self.assertIn("1 to 3", str(ctx.exception))
+
+ def test_each_array_type_builds_its_own_message(self):
+ cases = (
+ (dfb.double_array_column, common_pb2.DoubleArrayColumn, [[1.0, 2.0]]),
+ (dfb.float_array_column, common_pb2.FloatArrayColumn, [[1.5, 2.5]]),
+ (dfb.int32_array_column, common_pb2.Int32ArrayColumn, [[1, 2]]),
+ (dfb.int64_array_column, common_pb2.Int64ArrayColumn, [[1, 2]]),
+ (dfb.bool_array_column, common_pb2.BoolArrayColumn, [[True, False]]),
+ )
+ for builder, message_type, samples in cases:
+ with self.subTest(builder=builder.__name__):
+ column = builder("c", samples)
+ self.assertIsInstance(column, message_type)
+ self.assertEqual(list(column.values), samples[0])
+
+ def test_rejects_a_ragged_sample(self):
+ with self.assertRaises(ValueError) as ctx:
+ dfb.double_array_column("wf", [[[1.0, 2.0], [3.0]]])
+ message = str(ctx.exception)
+ self.assertIn("sample 0", message)
+ self.assertIn("rectangular", message)
+
+ def test_rejects_samples_of_different_shapes(self):
+ with self.assertRaises(ValueError) as ctx:
+ dfb.double_array_column("wf", [[1.0, 2.0], [1.0, 2.0, 3.0]])
+ message = str(ctx.exception)
+ self.assertIn("sample 1", message)
+ self.assertIn("same shape", message)
+
+ def test_rejects_scalar_samples(self):
+ with self.assertRaises(ValueError) as ctx:
+ dfb.double_array_column("wf", [1.0, 2.0])
+ self.assertIn("scalar samples", str(ctx.exception))
+
+ def test_rejects_no_samples_and_empty_samples(self):
+ with self.assertRaises(ValueError):
+ dfb.double_array_column("wf", [])
+ with self.assertRaises(ValueError) as ctx:
+ dfb.double_array_column("wf", [[], []])
+ self.assertIn("> 0", str(ctx.exception))
+
+ def test_rejects_blank_name(self):
+ with self.assertRaises(ValueError):
+ dfb.double_array_column(" ", [[1.0]])
+
+ def test_explicit_dims_shape_flat_samples(self):
+ column = dfb.double_array_column("img", [[1.0, 2.0, 3.0, 4.0]], dims=[2, 2])
+ self.assertEqual(list(column.dimensions.dims), [2, 2])
+ self.assertEqual(list(column.values), [1.0, 2.0, 3.0, 4.0])
+
+ def test_explicit_dims_may_restate_the_shape(self):
+ column = dfb.double_array_column("img", [[[1.0, 2.0], [3.0, 4.0]]], dims=[2, 2])
+ self.assertEqual(list(column.dimensions.dims), [2, 2])
+
+ def test_explicit_dims_that_do_not_describe_the_samples_are_rejected(self):
+ for samples, dims in (
+ ([[1.0, 2.0, 3.0, 4.0]], [3]), # wrong size for flat samples
+ ([[[1.0, 2.0, 3.0], [4.0, 5.0, 6.0]]], [3, 2]), # would reinterpret a 2x3 layout
+ ):
+ with self.subTest(dims=dims):
+ with self.assertRaises(ValueError) as ctx:
+ dfb.double_array_column("img", samples, dims=dims)
+ self.assertIn("does not describe", str(ctx.exception))
+
+ def test_rejects_a_value_the_type_cannot_hold(self):
+ for samples in ([[1.5]], [[2**40]]):
+ with self.subTest(samples=samples):
+ with self.assertRaises(ValueError) as ctx:
+ dfb.int32_array_column("i", samples)
+ self.assertIn("column 'i'", str(ctx.exception))
+
+ def test_metadata_is_carried(self):
+ metadata = dfb.column_metadata(tags=["raw"])
+ column = dfb.double_array_column("wf", [[1.0]], metadata=metadata)
+ self.assertEqual(list(column.metadata.tags), ["raw"])
+
+ def test_built_column_passes_data_frame(self):
+ column = dfb.double_array_column("wf", [[1.0, 2.0], [3.0, 4.0]])
+ frame = dfb.data_frame(_axis(2), [column])
+ self.assertEqual(len(frame.doubleArrayColumns), 1)
+
+ @unittest.skipUnless(_HAVE_NUMPY, "numpy not installed ([analysis] extra)")
+ def test_a_two_dimensional_numpy_array_is_one_sample_per_row(self):
+ samples = np.arange(6, dtype=np.float64).reshape(2, 3)
+ column = dfb.double_array_column("wf", samples)
+ self.assertEqual(list(column.dimensions.dims), [3])
+ self.assertEqual(list(column.values), [0.0, 1.0, 2.0, 3.0, 4.0, 5.0])
+
+ @unittest.skipUnless(_HAVE_NUMPY, "numpy not installed ([analysis] extra)")
+ def test_numpy_samples_keep_their_shape_and_flatten_row_major(self):
+ samples = [np.array([[1, 2, 3], [4, 5, 6]], dtype=np.int32)]
+ column = dfb.int32_array_column("m", samples)
+ self.assertEqual(list(column.dimensions.dims), [2, 3])
+ self.assertEqual(list(column.values), [1, 2, 3, 4, 5, 6])
+
+ @unittest.skipUnless(_HAVE_NUMPY, "numpy not installed ([analysis] extra)")
+ def test_numpy_bools_are_accepted(self):
+ column = dfb.bool_array_column("flags", [np.array([True, False])])
+ self.assertEqual(list(column.values), [True, False])
+
+ @unittest.skipUnless(_HAVE_NUMPY, "numpy not installed ([analysis] extra)")
+ def test_numpy_samples_of_different_shapes_are_rejected(self):
+ with self.assertRaises(ValueError) as ctx:
+ dfb.double_array_column("wf", [np.zeros(2), np.zeros(3)])
+ self.assertIn("same shape", str(ctx.exception))
+
+
+class TestImageStructSerializedBuilders(unittest.TestCase):
+ def test_image_column(self):
+ column = dfb.image_column("cam", [b"a", b"b"], width=4, height=3, channels=1, encoding="png")
+ descriptor = column.imageDescriptor
+ self.assertEqual(
+ (descriptor.width, descriptor.height, descriptor.channels, descriptor.encoding), (4, 3, 1, "png")
+ )
+ self.assertEqual(list(column.images), [b"a", b"b"])
+ dfb.data_frame(_axis(2), [column])
+
+ def test_image_column_rejections(self):
+ good = {"width": 4, "height": 3, "channels": 1, "encoding": "png"}
+ for label, overrides, images in (
+ ("no images", {}, []),
+ ("zero width", {"width": 0}, [b"a"]),
+ ("zero height", {"height": 0}, [b"a"]),
+ ("zero channels", {"channels": 0}, [b"a"]),
+ ("blank encoding", {"encoding": " "}, [b"a"]),
+ ):
+ with self.subTest(case=label), self.assertRaises(ValueError):
+ dfb.image_column("cam", images, **{**good, **overrides})
+
+ def test_struct_column(self):
+ column = dfb.struct_column("bpm", [b"s0", b"s1"], schema_id="beam_position:v3")
+ self.assertEqual(column.schemaId, "beam_position:v3")
+ self.assertEqual(list(column.values), [b"s0", b"s1"])
+ dfb.data_frame(_axis(2), [column])
+
+ def test_struct_column_rejections(self):
+ with self.assertRaises(ValueError):
+ dfb.struct_column("bpm", [b"s0"], schema_id=" ")
+ with self.assertRaises(ValueError):
+ dfb.struct_column("bpm", [], schema_id="s")
+
+ def test_serialized_column(self):
+ column = dfb.serialized_column("blob", b"payload", encoding="proto:Image")
+ self.assertEqual((column.encoding, column.payload), ("proto:Image", b"payload"))
+ dfb.data_frame(_axis(3), [dfb.double_column("d", [1.0, 2.0, 3.0]), column])
+
+ def test_serialized_column_rejections(self):
+ with self.assertRaises(ValueError):
+ dfb.serialized_column("blob", b"p", encoding="")
+ with self.assertRaises(ValueError):
+ dfb.serialized_column("", b"p", encoding="e")
+
+
+# A present-day start with a nanosecond fraction, and a period that does not divide a second: a float anywhere in
+# the chunk start-time arithmetic would show up here.
+_ODD_START = common_pb2.Timestamp(epochSeconds=int(T0.timestamp()), nanoseconds=123_456_789)
+_ODD_PERIOD = 333_333_333
+
+
+def _nanos(timestamp):
+ return timestamp.epochSeconds * 1_000_000_000 + timestamp.nanoseconds
+
+
+def _every_kind_frame(count):
+ """A frame carrying one column of every splittable kind, each with metadata, on an odd SamplingClock."""
+ metadata = dfb.column_metadata(tags=["t"], attributes={"k": "v"})
+ columns = [
+ dfb.double_column("double", [float(i) for i in range(count)], metadata=metadata),
+ dfb.float_column("float", [i + 0.5 for i in range(count)]),
+ dfb.int64_column("int64", list(range(count))),
+ dfb.int32_column("int32", list(range(count))),
+ dfb.bool_column("bool", [i % 2 == 0 for i in range(count)]),
+ dfb.string_column("string", [f"s{i}" for i in range(count)]),
+ dfb.enum_column("enum", [i % 3 for i in range(count)], enum_id="E"),
+ dfb.struct_column("struct", [bytes([i]) for i in range(count)], schema_id="S:v1"),
+ dfb.image_column("image", [bytes([i, i]) for i in range(count)], 2, 1, 1, "raw-u8"),
+ dfb.double_array_column("darr", [[i, i + 0.5] for i in range(count)], metadata=metadata),
+ dfb.float_array_column("farr", [[[float(i)], [float(i)]] for i in range(count)]),
+ dfb.int32_array_column("iarr", [[i, -i] for i in range(count)]),
+ dfb.int64_array_column("larr", [[i] for i in range(count)]),
+ dfb.bool_array_column("barr", [[True, i % 2 == 0] for i in range(count)]),
+ dfb.data_column("legacy", [i if i % 2 else None for i in range(count)]),
+ ]
+ return dfb.data_frame(dfb.sampling_clock(_ODD_START, _ODD_PERIOD, count), columns)
+
+
+def _assert_chunks_reproduce(test, frame, chunks):
+ """Concatenating the chunks' timestamps and per-column values gives back the frame's."""
+ from dp_python_lib.client import data_frame_conversions as dfc
+
+ test.assertEqual(
+ [t for chunk in chunks for t in dfc.data_frame_timestamps(chunk)], dfc.data_frame_timestamps(frame)
+ )
+ expected = dfc.data_frame_columns(frame)
+ combined = {name: [] for name in expected}
+ for chunk in chunks:
+ test.assertEqual(set(dfc.data_frame_columns(chunk)), set(expected))
+ for name, values in dfc.data_frame_columns(chunk).items():
+ combined[name].extend(values)
+ test.assertEqual(combined, expected)
+
+
+class TestSplitDataFrame(unittest.TestCase):
+ def test_requires_a_limit(self):
+ with self.assertRaises(ValueError) as ctx:
+ dfb.split_data_frame(_every_kind_frame(3))
+ self.assertIn("at least one of", str(ctx.exception))
+
+ def test_rejects_out_of_range_limits(self):
+ frame = _every_kind_frame(3)
+ for kwargs in ({"max_rows": 0}, {"max_bytes": 0}, {"max_span_nanos": -1}):
+ with self.subTest(**kwargs), self.assertRaises(ValueError):
+ dfb.split_data_frame(frame, **kwargs)
+
+ def test_validates_eagerly_not_on_first_iteration(self):
+ frame = common_pb2.DataFrame()
+ frame.dataTimestamps.CopyFrom(_axis(2))
+ frame.enumColumns.add(name="e").values[:] = [0, 1] # no enumId
+ with self.assertRaises(ValueError):
+ dfb.split_data_frame(frame, max_rows=1)
+
+ def test_rejects_a_frame_with_a_serialized_column(self):
+ frame = dfb.data_frame(_axis(2), [dfb.double_column("d", [1.0, 2.0]), dfb.serialized_column("s", b"x", "e")])
+ with self.assertRaises(ValueError) as ctx:
+ dfb.split_data_frame(frame, max_rows=1)
+ self.assertIn("serialized", str(ctx.exception))
+
+ def test_is_lazy(self):
+ chunks = dfb.split_data_frame(_every_kind_frame(10), max_rows=3)
+ self.assertNotIsInstance(chunks, (list, tuple))
+ self.assertEqual(next(chunks).dataTimestamps.samplingClock.count, 3)
+
+ def test_max_rows_splits_every_column_kind_and_reproduces_the_frame(self):
+ frame = _every_kind_frame(10)
+ chunks = list(dfb.split_data_frame(frame, max_rows=4))
+ self.assertEqual([c.dataTimestamps.samplingClock.count for c in chunks], [4, 4, 2])
+ for chunk in chunks:
+ dfb.validate_data_frame(chunk)
+ _assert_chunks_reproduce(self, frame, chunks)
+
+ def test_sampling_clock_chunk_starts_are_exact_integer_nanoseconds(self):
+ frame = _every_kind_frame(10)
+ chunks = list(dfb.split_data_frame(frame, max_rows=3))
+ starts = [_nanos(c.dataTimestamps.samplingClock.startTime) for c in chunks]
+ self.assertEqual(starts, [_nanos(_ODD_START) + row * _ODD_PERIOD for row in (0, 3, 6, 9)])
+ for chunk in chunks:
+ self.assertEqual(chunk.dataTimestamps.samplingClock.periodNanos, _ODD_PERIOD)
+
+ def test_structural_fields_and_metadata_travel_with_every_chunk(self):
+ for chunk in dfb.split_data_frame(_every_kind_frame(5), max_rows=2):
+ self.assertEqual(chunk.enumColumns[0].enumId, "E")
+ self.assertEqual(chunk.structColumns[0].schemaId, "S:v1")
+ self.assertEqual(chunk.imageColumns[0].imageDescriptor.encoding, "raw-u8")
+ self.assertEqual(list(chunk.floatArrayColumns[0].dimensions.dims), [2, 1])
+ self.assertEqual(list(chunk.doubleColumns[0].metadata.tags), ["t"])
+ self.assertEqual(list(chunk.doubleArrayColumns[0].metadata.tags), ["t"])
+
+ def test_timestamp_list_is_sliced(self):
+ axis = dfb.timestamp_list([100, 101, 105, 200, 201])
+ frame = dfb.data_frame(axis, [dfb.double_column("d", [1.0, 2.0, 3.0, 4.0, 5.0])])
+ chunks = list(dfb.split_data_frame(frame, max_rows=2))
+ self.assertEqual(
+ [[t.epochSeconds for t in c.dataTimestamps.timestampList.timestamps] for c in chunks],
+ [[100, 101], [105, 200], [201]],
+ )
+ _assert_chunks_reproduce(self, frame, chunks)
+
+ def test_max_span_on_a_sampling_clock_allows_exactly_the_limit(self):
+ # The server rejects a span GREATER than its cap, measured first-to-last: (count - 1) * period.
+ frame = dfb.data_frame(_axis(10), [dfb.double_column("d", [float(i) for i in range(10)])])
+ chunks = list(dfb.split_data_frame(frame, max_span_nanos=3 * 1_000_000_000))
+ self.assertEqual([c.dataTimestamps.samplingClock.count for c in chunks], [4, 4, 2])
+
+ def test_max_span_on_a_timestamp_list(self):
+ axis = dfb.timestamp_list([100, 101, 105, 200, 201])
+ frame = dfb.data_frame(axis, [dfb.double_column("d", [1.0, 2.0, 3.0, 4.0, 5.0])])
+ chunks = list(dfb.split_data_frame(frame, max_span_nanos=5 * 1_000_000_000))
+ self.assertEqual(
+ [[t.epochSeconds for t in c.dataTimestamps.timestampList.timestamps] for c in chunks],
+ [[100, 101, 105], [200, 201]],
+ )
+
+ def test_zero_span_gives_one_row_per_chunk(self):
+ frame = dfb.data_frame(_axis(3), [dfb.double_column("d", [1.0, 2.0, 3.0])])
+ self.assertEqual(len(list(dfb.split_data_frame(frame, max_span_nanos=0))), 3)
+
+ def test_the_tightest_limit_wins(self):
+ frame = dfb.data_frame(_axis(10), [dfb.double_column("d", [float(i) for i in range(10)])])
+ chunks = list(dfb.split_data_frame(frame, max_rows=3, max_span_nanos=1_000_000_000))
+ self.assertEqual([c.dataTimestamps.samplingClock.count for c in chunks], [2, 2, 2, 2, 2])
+
+ def test_budgeted_request_size_matches_a_real_request_with_worst_case_ids(self):
+ # The envelope arithmetic must agree with protobuf's own measurement of the message the server receives.
+ worst_id = "\U0010ffff" * dfb.MAX_BUDGETED_ID_CHARS
+ for count in (1, 100, 20_000):
+ with self.subTest(count=count):
+ frame = dfb.data_frame(_axis(count), [dfb.double_column("d", [0.5] * count)])
+ request = ingestion_pb2.IngestDataRequest(
+ providerId=worst_id, clientRequestId=worst_id, ingestionDataFrame=frame
+ )
+ self.assertEqual(dfb._budgeted_request_bytes(frame.ByteSize()), request.ByteSize())
+
+ def test_every_chunk_fits_max_bytes_as_a_request_with_worst_case_ids(self):
+ worst_id = "\U0010ffff" * dfb.MAX_BUDGETED_ID_CHARS
+ frame = _every_kind_frame(200)
+ max_bytes = 6_000
+ chunks = list(dfb.split_data_frame(frame, max_bytes=max_bytes))
+ self.assertGreater(len(chunks), 1)
+ for chunk in chunks:
+ request = ingestion_pb2.IngestDataRequest(
+ providerId=worst_id, clientRequestId=worst_id, ingestionDataFrame=chunk
+ )
+ self.assertLessEqual(request.ByteSize(), max_bytes)
+ _assert_chunks_reproduce(self, frame, chunks)
+
+ def test_max_bytes_holds_when_row_sizes_vary(self):
+ # Rows early in the frame are tiny and later ones large, so a size estimated from the first rows overshoots.
+ values = ["x"] * 50 + ["y" * 200] * 50
+ frame = dfb.data_frame(_axis(100), [dfb.string_column("s", values)])
+ max_bytes = 4_000
+ chunks = list(dfb.split_data_frame(frame, max_bytes=max_bytes))
+ for chunk in chunks:
+ self.assertLessEqual(dfb._budgeted_request_bytes(chunk.ByteSize()), max_bytes)
+ _assert_chunks_reproduce(self, frame, chunks)
+
+ def test_a_row_too_large_for_max_bytes_raises_naming_it(self):
+ frame = dfb.data_frame(_axis(3), [dfb.string_column("s", ["a", "b" * 5_000, "c"])])
+ chunks = dfb.split_data_frame(frame, max_bytes=4_000)
+ self.assertEqual(next(chunks).dataTimestamps.samplingClock.count, 1)
+ with self.assertRaises(ValueError) as ctx:
+ next(chunks)
+ self.assertIn("row 1", str(ctx.exception))
+
+ def test_server_default_is_exported_as_a_reference(self):
+ self.assertEqual(dfb.SERVER_DEFAULT_MAX_MESSAGE_BYTES, 4_096_000)
if __name__ == "__main__":
diff --git a/tests/unit/test_data_frame_conversions.py b/tests/unit/test_data_frame_conversions.py
index e1f21c2..a34c19d 100644
--- a/tests/unit/test_data_frame_conversions.py
+++ b/tests/unit/test_data_frame_conversions.py
@@ -824,5 +824,61 @@ def test_other_column_kinds_have_neither(self):
self.assertEqual(dfc.data_frame_schema_ids(frame), {})
+class TestBuilderReadBackSymmetry(unittest.TestCase):
+ """
+ Each non-scalar builder (#17, D7) round-trips through data_frame() and the read side: values, one entry per
+ sample, plus the structural field the payload needs. Live read-back waits for the bucket query (#16), so this
+ is the check that the write and read halves agree on layout -- above all, row-major array flattening.
+ """
+
+ def test_array_builders_round_trip_values_and_dims(self):
+ cases = (
+ (dfb.double_array_column, [[[1.0, 2.0, 3.0], [4.0, 5.0, 6.0]], [[7.0, 8.0, 9.0], [10.0, 11.0, 12.0]]]),
+ (dfb.float_array_column, [[0.5, 1.5], [2.5, 3.5]]),
+ (dfb.int32_array_column, [[[1], [2]], [[3], [4]]]),
+ (dfb.int64_array_column, [[2**40, -1], [0, 1]]),
+ (dfb.bool_array_column, [[True, False, True], [False, False, True]]),
+ )
+ for builder, samples in cases:
+ with self.subTest(builder=builder.__name__):
+ frame = dfb.data_frame(_axis(2), [builder("arr", samples)])
+ expected_dims = [len(samples[0])]
+ if isinstance(samples[0][0], list):
+ expected_dims.append(len(samples[0][0]))
+ self.assertEqual(dfc.data_frame_column_dimensions(frame), {"arr": expected_dims})
+ # column_values() keeps each sample flat; flattening the input row-major must reproduce it.
+ flat = [
+ [v for row in sample for v in row] if isinstance(sample[0], list) else sample for sample in samples
+ ]
+ self.assertEqual(dfc.data_frame_columns(frame), {"arr": flat})
+
+ def test_explicit_dims_round_trip(self):
+ frame = dfb.data_frame(_axis(1), [dfb.double_array_column("img", [[1.0, 2.0, 3.0, 4.0]], dims=[2, 2])])
+ self.assertEqual(dfc.data_frame_column_dimensions(frame), {"img": [2, 2]})
+ self.assertEqual(dfc.data_frame_columns(frame), {"img": [[1.0, 2.0, 3.0, 4.0]]})
+
+ def test_image_builder_round_trips_payloads_and_descriptor(self):
+ column = dfb.image_column("cam", [b"f0", b"f1"], width=640, height=480, channels=3, encoding="rgb8")
+ frame = dfb.data_frame(_axis(2), [column])
+ self.assertEqual(dfc.data_frame_columns(frame), {"cam": [b"f0", b"f1"]})
+ self.assertEqual(
+ dfc.data_frame_image_descriptors(frame),
+ {"cam": {"width": 640, "height": 480, "channels": 3, "encoding": "rgb8"}},
+ )
+
+ def test_struct_builder_round_trips_payloads_and_schema_id(self):
+ frame = dfb.data_frame(_axis(2), [dfb.struct_column("bpm", [b"s0", b"s1"], schema_id="beam_position:v3")])
+ self.assertEqual(dfc.data_frame_columns(frame), {"bpm": [b"s0", b"s1"]})
+ self.assertEqual(dfc.data_frame_schema_ids(frame), {"bpm": "beam_position:v3"})
+
+ def test_serialized_builder_is_carried_but_not_decoded(self):
+ # The read side skips serialized columns deliberately (their payload is opaque); the message keeps them.
+ frame = dfb.data_frame(
+ _axis(1), [dfb.double_column("d", [1.0]), dfb.serialized_column("blob", b"payload", "proto:Image")]
+ )
+ self.assertEqual(dfc.data_frame_columns(frame), {"d": [1.0]})
+ self.assertEqual([(c.name, c.payload) for c in frame.serializedDataColumns], [("blob", b"payload")])
+
+
if __name__ == "__main__":
unittest.main()
diff --git a/tests/unit/test_ingestion_client.py b/tests/unit/test_ingestion_client.py
index c94d5a9..56ac010 100644
--- a/tests/unit/test_ingestion_client.py
+++ b/tests/unit/test_ingestion_client.py
@@ -1,18 +1,27 @@
+import contextlib
import os
import sys
import unittest
-from unittest.mock import Mock
+from datetime import datetime, timezone
+from unittest.mock import Mock, patch
import grpc
# Add src directory to path for imports
sys.path.insert(0, os.path.join(os.path.dirname(__file__), "../../src"))
+from dp_python_lib.client import data_frame as dfb
from dp_python_lib.client.ingestion_client import (
+ REQUEST_STATUS_CLOCK_SKEW,
+ IngestDataRequestParams,
IngestionClient,
+ IngestionRequestStatus,
RegisterProviderApiResult,
RegisterProviderRequestParams,
+ chunked_request_params,
)
+from dp_python_lib.client.ingestion_client import RequestStatusQuery as RS
+from dp_python_lib.client.time_conversions import to_epoch_nanos
from dp_python_lib.grpc import common_pb2, ingestion_pb2
@@ -267,5 +276,555 @@ def test_send_register_provider_general_exception(self):
self.assertIsNone(result.response)
+# ----------------------------------------------------------------------
+# #17: data ingestion and request status
+# ----------------------------------------------------------------------
+
+PROVIDER = "provider-1"
+SINCE = datetime(2026, 9, 30, 12, 0, 0, tzinfo=timezone.utc)
+_RequestStatus = ingestion_pb2.QueryRequestStatusResponse.RequestStatusResult.RequestStatus
+
+
+def _frame(count=2):
+ return dfb.data_frame(
+ dfb.sampling_clock(1_770_000_000, 1_000, count), [dfb.double_column("pv", [float(i) for i in range(count)])]
+ )
+
+
+def _params(request_id="req-1"):
+ return IngestDataRequestParams(PROVIDER, _frame(), client_request_id=request_id)
+
+
+class _FakeRpcError(grpc.RpcError):
+ """An RpcError with a resolvable code(), unlike the bare grpc.RpcError() the older tests raise."""
+
+ def __init__(self, code, details):
+ super().__init__()
+ self._code, self._details = code, details
+
+ def code(self):
+ return self._code
+
+ def details(self):
+ return self._details
+
+
+def _ack(request_id, rows=2, columns=1):
+ response = ingestion_pb2.IngestDataResponse(providerId=PROVIDER, clientRequestId=request_id)
+ response.ackResult.numRows, response.ackResult.numColumns = rows, columns
+ return response
+
+
+def _reject(request_id, message="bad request"):
+ response = ingestion_pb2.IngestDataResponse(providerId=PROVIDER, clientRequestId=request_id)
+ response.exceptionalResult.message = message
+ return response
+
+
+def _status_response(*documents):
+ response = ingestion_pb2.QueryRequestStatusResponse()
+ response.requestStatusResult.requestStatus.extend(documents)
+ return response
+
+
+def _document(request_id, status=IngestionRequestStatus.SUCCESS):
+ return _RequestStatus(providerId=PROVIDER, requestId=request_id, ingestionRequestStatus=int(status))
+
+
+class _ClientTestCase(unittest.TestCase):
+ def setUp(self):
+ self.client = IngestionClient(Mock())
+ self.stub = Mock()
+ self.client._stub = self.stub
+
+
+class TestRegisterProviderResultProperties(unittest.TestCase):
+ def test_success_exposes_id_and_newness(self):
+ response = ingestion_pb2.RegisterProviderResponse()
+ response.registrationResult.providerId = "p-1"
+ response.registrationResult.isNewProvider = True
+ result = RegisterProviderApiResult(is_error=False, message="", response=response)
+ self.assertEqual((result.provider_id, result.is_new_provider), ("p-1", True))
+
+ def test_error_has_neither(self):
+ result = RegisterProviderApiResult(is_error=True, message="nope")
+ self.assertIsNone(result.provider_id)
+ self.assertIsNone(result.is_new_provider)
+
+
+class TestIngestDataRequestParams(unittest.TestCase):
+ def test_generates_a_uuid_request_id_by_default(self):
+ first, second = IngestDataRequestParams(PROVIDER, _frame()), IngestDataRequestParams(PROVIDER, _frame())
+ self.assertEqual(len(first.client_request_id), 36)
+ self.assertNotEqual(first.client_request_id, second.client_request_id)
+
+ def test_keeps_an_explicit_request_id(self):
+ self.assertEqual(_params("mine").client_request_id, "mine")
+
+ def test_rejects_blank_ids(self):
+ for kwargs in (
+ {"provider_id": " "},
+ {"provider_id": ""},
+ {"client_request_id": "\t"},
+ {"client_request_id": ""},
+ ):
+ with self.subTest(**kwargs), self.assertRaises(ValueError):
+ IngestDataRequestParams(**{"provider_id": PROVIDER, "frame": _frame(), **kwargs})
+
+ def test_id_length_is_capped_at_the_chunking_budget(self):
+ # split_data_frame() reserves room for ids of MAX_BUDGETED_ID_CHARS; a longer one could overrun max_bytes.
+ IngestDataRequestParams(PROVIDER, _frame(), client_request_id="x" * dfb.MAX_BUDGETED_ID_CHARS)
+ for kwargs in ({"client_request_id": "x" * 257}, {"provider_id": "p" * 257}):
+ with self.subTest(field=next(iter(kwargs))), self.assertRaises(ValueError) as ctx:
+ IngestDataRequestParams(**{"provider_id": PROVIDER, "frame": _frame(), **kwargs})
+ self.assertIn("at most 256", str(ctx.exception))
+
+ def test_frame_is_revalidated(self):
+ frame = common_pb2.DataFrame()
+ frame.dataTimestamps.CopyFrom(dfb.sampling_clock(1_770_000_000, 1_000, 1))
+ frame.enumColumns.add(name="state").values[:] = [0] # hand-built, no enumId
+ with self.assertRaises(ValueError) as ctx:
+ IngestDataRequestParams(PROVIDER, frame)
+ self.assertIn("enumId", str(ctx.exception))
+ # Named for what the caller called, not data_frame(), which this frame never went through.
+ self.assertTrue(str(ctx.exception).startswith("IngestDataRequestParams "), str(ctx.exception))
+
+
+class TestChunkedRequestParams(unittest.TestCase):
+ def test_numbers_the_chunks_under_one_base(self):
+ chunks = dfb.split_data_frame(_frame(5), max_rows=2)
+ params = list(chunked_request_params(PROVIDER, chunks, base_request_id="load-7"))
+ self.assertEqual([p.client_request_id for p in params], ["load-7-0", "load-7-1", "load-7-2"])
+ self.assertTrue(all(p.provider_id == PROVIDER for p in params))
+
+ def test_default_base_is_generated_and_shared(self):
+ params = list(chunked_request_params(PROVIDER, [_frame(), _frame()]))
+ bases = {p.client_request_id.rsplit("-", 1)[0] for p in params}
+ self.assertEqual(len(bases), 1)
+
+ def test_is_lazy(self):
+ def frames():
+ yield _frame()
+ raise AssertionError("consumed too far")
+
+ self.assertEqual(next(chunked_request_params(PROVIDER, frames())).client_request_id[-2:], "-0")
+
+ def test_bad_base_fails_as_itself_before_any_frame_is_consumed(self):
+ def frames():
+ raise AssertionError("a frame was consumed")
+ yield
+
+ for base, expected in (("", "non-blank"), (" ", "non-blank"), ("b" * 255, "base_request_id is 255")):
+ with self.subTest(base=base[:5]), self.assertRaises(ValueError) as ctx:
+ next(chunked_request_params(PROVIDER, frames(), base_request_id=base))
+ self.assertIn("base_request_id", str(ctx.exception))
+ self.assertIn(expected, str(ctx.exception))
+
+ def test_longest_base_that_fits_the_first_suffix_is_accepted(self):
+ (params,) = chunked_request_params(PROVIDER, [_frame()], base_request_id="b" * 254)
+ self.assertEqual(len(params.client_request_id), dfb.MAX_BUDGETED_ID_CHARS)
+
+
+class TestIngestData(_ClientTestCase):
+ def test_builds_the_request(self):
+ params = _params("r-1")
+ request = self.client._build_ingest_data_request(params)
+ self.assertEqual((request.providerId, request.clientRequestId), (PROVIDER, "r-1"))
+ self.assertEqual(request.ingestionDataFrame, params.frame)
+
+ def test_ack(self):
+ self.stub.ingestData.return_value = _ack("r-1")
+ result = self.client.ingest_data(_params("r-1"))
+ self.assertFalse(result.result_status.is_error)
+ self.assertEqual((result.provider_id, result.client_request_id), (PROVIDER, "r-1"))
+ self.assertEqual((result.num_rows, result.num_columns), (2, 1))
+ sent = self.stub.ingestData.call_args[0][0]
+ self.assertEqual(sent.clientRequestId, "r-1")
+
+ def test_reject_still_names_the_request(self):
+ self.stub.ingestData.return_value = _reject("r-1", "bad frame")
+ result = self.client.ingest_data(_params("r-1"))
+ self.assertTrue(result.result_status.is_error)
+ self.assertEqual(result.result_status.message, "bad frame")
+ self.assertEqual((result.provider_id, result.client_request_id), (PROVIDER, "r-1"))
+ self.assertIsNone(result.num_rows)
+
+ def test_unrecognized_response_is_an_error(self):
+ self.stub.ingestData.return_value = ingestion_pb2.IngestDataResponse(clientRequestId="r-1")
+ result = self.client.ingest_data(_params("r-1"))
+ self.assertIn("neither exceptionalResult nor ackResult", result.result_status.message)
+
+ def test_grpc_error(self):
+ self.stub.ingestData.side_effect = _FakeRpcError(grpc.StatusCode.UNAVAILABLE, "connection refused")
+ result = self.client.ingest_data(_params())
+ self.assertEqual(result.result_status.message, "gRPC error: connection refused")
+
+ def test_unexpected_exception(self):
+ self.stub.ingestData.side_effect = RuntimeError("kaboom")
+ result = self.client.ingest_data(_params())
+ self.assertIn("Unexpected error: kaboom", result.result_status.message)
+
+ def test_oversized_request_points_at_split_data_frame(self):
+ self.stub.ingestData.side_effect = _FakeRpcError(
+ grpc.StatusCode.RESOURCE_EXHAUSTED, "gRPC message exceeds maximum size 4096000: 5000000"
+ )
+ message = self.client.ingest_data(_params()).result_status.message
+ self.assertTrue(message.startswith("gRPC error: gRPC message exceeds maximum size"))
+ self.assertIn("split_data_frame", message)
+
+ def test_other_resource_exhaustion_gets_no_hint(self):
+ for code, details in (
+ (grpc.StatusCode.RESOURCE_EXHAUSTED, "quota exceeded"),
+ (grpc.StatusCode.UNKNOWN, "gRPC message exceeds maximum size"),
+ ):
+ with self.subTest(code=code):
+ self.stub.ingestData.side_effect = _FakeRpcError(code, details)
+ self.assertEqual(self.client.ingest_data(_params()).result_status.message, f"gRPC error: {details}")
+
+
+def _consume_then(response):
+ """A stub side effect that drains the request iterator, as grpcio does, before returning response."""
+ sent = []
+
+ def call(requests):
+ sent.extend(requests)
+ return response
+
+ return call, sent
+
+
+class TestIngestDataStream(_ClientTestCase):
+ def test_all_accepted(self):
+ response = ingestion_pb2.IngestDataStreamResponse(clientRequestIds=["a", "b"])
+ response.ingestDataStreamResult.numRequests = 2
+ self.stub.ingestDataStream.side_effect, sent = _consume_then(response)
+ result = self.client.ingest_data_stream([_params("a"), _params("b")])
+ self.assertFalse(result.result_status.is_error)
+ self.assertEqual(result.num_requests, 2)
+ self.assertEqual([r.clientRequestId for r in sent], ["a", "b"])
+
+ def test_partial_reject_keeps_the_response(self):
+ response = ingestion_pb2.IngestDataStreamResponse(clientRequestIds=["a", "b"], rejectedRequestIds=["b"])
+ response.exceptionalResult.message = "one or more requests were rejected"
+ self.stub.ingestDataStream.side_effect, _ = _consume_then(response)
+ result = self.client.ingest_data_stream([_params("a"), _params("b")])
+ self.assertTrue(result.result_status.is_error)
+ self.assertEqual(result.rejected_request_ids, ["b"])
+ self.assertEqual(result.client_request_ids, ["a", "b"])
+ self.assertIsNone(result.num_requests)
+
+ def test_grpc_error_has_no_response(self):
+ self.stub.ingestDataStream.side_effect = _FakeRpcError(grpc.StatusCode.UNAVAILABLE, "down")
+ result = self.client.ingest_data_stream([_params()])
+ self.assertEqual(result.result_status.message, "gRPC error: down")
+ self.assertIsNone(result.client_request_ids)
+ self.assertIsNone(result.rejected_request_ids)
+ self.assertIsNone(result.num_requests)
+
+ def test_oversized_request_points_at_split_data_frame(self):
+ self.stub.ingestDataStream.side_effect = _FakeRpcError(
+ grpc.StatusCode.RESOURCE_EXHAUSTED, "gRPC message exceeds maximum size 4096000: 9"
+ )
+ self.assertIn("split_data_frame", self.client.ingest_data_stream([_params()]).result_status.message)
+
+ def test_unexpected_exception_has_no_response(self):
+ self.stub.ingestDataStream.side_effect = RuntimeError("kaboom")
+ result = self.client.ingest_data_stream([_params()])
+ self.assertTrue(result.result_status.is_error)
+ self.assertEqual(result.result_status.message, "Unexpected error: kaboom")
+ self.assertIsNone(result.response)
+ self.assertIsNone(result.client_request_ids)
+ self.assertIsNone(result.rejected_request_ids)
+ self.assertIsNone(result.num_requests)
+
+ def test_unrecognized_response_is_an_error(self):
+ self.stub.ingestDataStream.side_effect, _ = _consume_then(ingestion_pb2.IngestDataStreamResponse())
+ result = self.client.ingest_data_stream([_params()])
+ self.assertIn("neither exceptionalResult nor ingestDataStreamResult", result.result_status.message)
+
+
+class TestIngestDataBidiStream(_ClientTestCase):
+ def test_yields_rejects_without_raising(self):
+ self.stub.ingestDataBidiStream.return_value = iter([_ack("a"), _reject("b", "nope"), _ack("c")])
+ results = list(self.client.iter_ingest_data_bidi_stream([_params("a"), _params("b"), _params("c")]))
+ self.assertEqual([r.client_request_id for r in results], ["a", "b", "c"])
+ self.assertEqual([r.result_status.is_error for r in results], [False, True, False])
+ self.assertEqual(results[1].result_status.message, "nope")
+
+ def test_unrecognized_response_is_yielded_as_that_requests_error(self):
+ self.stub.ingestDataBidiStream.return_value = iter([ingestion_pb2.IngestDataResponse(clientRequestId="a")])
+ (result,) = list(self.client.iter_ingest_data_bidi_stream([_params("a")]))
+ self.assertTrue(result.result_status.is_error)
+ self.assertEqual(result.client_request_id, "a")
+
+ def test_transport_error_mid_stream_raises_after_earlier_results(self):
+ def responses():
+ yield _ack("a")
+ raise _FakeRpcError(grpc.StatusCode.UNAVAILABLE, "stream reset")
+
+ self.stub.ingestDataBidiStream.return_value = responses()
+ seen = []
+ with self.assertRaises(RuntimeError) as ctx:
+ for result in self.client.iter_ingest_data_bidi_stream([_params("a"), _params("b")]):
+ seen.append(result.client_request_id)
+ self.assertEqual(seen, ["a"])
+ self.assertIn("gRPC error: stream reset", str(ctx.exception))
+
+ def test_unexpected_exception_raises_runtime_error(self):
+ self.stub.ingestDataBidiStream.side_effect = RuntimeError("kaboom")
+ with self.assertRaises(RuntimeError) as ctx:
+ list(self.client.iter_ingest_data_bidi_stream([_params("a")]))
+ self.assertEqual(str(ctx.exception), "ingestDataBidiStream failed: Unexpected error: kaboom")
+
+ def test_unexpected_exception_mid_stream_raises_after_earlier_results(self):
+ def responses():
+ yield _ack("a")
+ raise RuntimeError("kaboom")
+
+ self.stub.ingestDataBidiStream.return_value = responses()
+ seen = []
+ with self.assertRaises(RuntimeError) as ctx:
+ for result in self.client.iter_ingest_data_bidi_stream([_params("a"), _params("b")]):
+ seen.append(result.client_request_id)
+ self.assertEqual(seen, ["a"])
+ self.assertIn("Unexpected error: kaboom", str(ctx.exception))
+
+
+class _CancellableResponses:
+ """A response iterator with the cancel() a grpcio stream-stream call has."""
+
+ def __init__(self, responses):
+ self._responses = iter(responses)
+ self.cancel = Mock()
+
+ def __iter__(self):
+ return self._responses
+
+
+class TestIngestDataBidiStreamCancellation(_ClientTestCase):
+ def test_stopping_early_cancels_the_call(self):
+ # Otherwise grpcio keeps sending requests from its own thread after the caller has walked away.
+ responses = _CancellableResponses([_ack("a"), _ack("b")])
+ self.stub.ingestDataBidiStream.return_value = responses
+ results = self.client.iter_ingest_data_bidi_stream([_params("a"), _params("b")])
+ next(results)
+ results.close()
+ responses.cancel.assert_called_once()
+
+ def test_closing_around_a_break_cancels_the_call(self):
+ # The documented way to stop early.
+ responses = _CancellableResponses([_ack("a"), _ack("b")])
+ self.stub.ingestDataBidiStream.return_value = responses
+ with contextlib.closing(self.client.iter_ingest_data_bidi_stream([_params("a"), _params("b")])) as results:
+ for _result in results:
+ break
+ responses.cancel.assert_called_once()
+
+ def test_a_bare_break_on_a_retained_iterator_does_not_cancel(self):
+ # Pins the premise behind the docs asking for an explicit close: breaking out of a loop leaves a generator
+ # that is still referenced suspended, so its finally -- and the cancel -- have not run.
+ responses = _CancellableResponses([_ack("a"), _ack("b")])
+ self.stub.ingestDataBidiStream.return_value = responses
+ results = self.client.iter_ingest_data_bidi_stream([_params("a"), _params("b")])
+ for _result in results:
+ break
+ responses.cancel.assert_not_called()
+ results.close()
+ responses.cancel.assert_called_once()
+
+ def test_a_completed_call_is_cancelled_harmlessly(self):
+ responses = _CancellableResponses([_ack("a")])
+ self.stub.ingestDataBidiStream.return_value = responses
+ list(self.client.iter_ingest_data_bidi_stream([_params("a")]))
+ responses.cancel.assert_called_once()
+
+
+class TestRequestStatusQuery(unittest.TestCase):
+ def test_enum_mirrors_the_proto(self):
+ self.assertEqual(
+ {member.name: member.value for member in IngestionRequestStatus},
+ {
+ name.replace("INGESTION_REQUEST_STATUS_", ""): value
+ for name, value in ingestion_pb2.IngestionRequestStatus.items()
+ },
+ )
+
+ def test_id_and_name_criteria(self):
+ self.assertEqual(RS.provider_id("p").providerIdCriterion.providerId, "p")
+ self.assertEqual(RS.provider_name("n").providerNameCriterion.providerName, "n")
+ self.assertEqual(RS.request_id("r").requestIdCriterion.requestId, "r")
+
+ def test_blank_inputs_are_rejected(self):
+ for helper in (RS.provider_id, RS.provider_name, RS.request_id):
+ with self.subTest(helper=helper.__name__), self.assertRaises(ValueError):
+ helper(" ")
+
+ def test_status_accepts_members_and_ints(self):
+ criterion = RS.status([IngestionRequestStatus.ERROR, 1])
+ self.assertEqual(list(criterion.statusCriterion.status), [2, 1])
+
+ def test_status_rejects_empty_and_unknown(self):
+ with self.assertRaises(ValueError):
+ RS.status([])
+ with self.assertRaises(ValueError) as ctx:
+ RS.status([7])
+ self.assertIn("unknown status", str(ctx.exception))
+
+ def test_time_range_with_and_without_end(self):
+ open_ended = RS.time_range(SINCE)
+ self.assertTrue(open_ended.timeRangeCriterion.HasField("beginTime"))
+ self.assertFalse(open_ended.timeRangeCriterion.HasField("endTime"))
+ closed = RS.time_range(SINCE, SINCE) # inclusive at both ends, so equal bounds are a valid instant
+ self.assertTrue(closed.timeRangeCriterion.HasField("endTime"))
+
+ def test_time_range_rejections(self):
+ with self.assertRaises(ValueError):
+ RS.time_range(0) # the server requires begin >= 1 s
+ with self.assertRaises(ValueError):
+ RS.time_range(SINCE, datetime(2026, 1, 1, tzinfo=timezone.utc))
+
+
+class TestQueryRequestStatus(_ClientTestCase):
+ def test_requires_a_criterion(self):
+ with self.assertRaises(ValueError):
+ self.client.query_request_status([])
+
+ def test_returns_the_documents(self):
+ self.stub.queryRequestStatus.return_value = _status_response(_document("a"), _document("b"))
+ result = self.client.query_request_status([RS.provider_id(PROVIDER)])
+ self.assertEqual([d.requestId for d in result.request_statuses], ["a", "b"])
+ sent = self.stub.queryRequestStatus.call_args[0][0]
+ self.assertEqual(sent.criteria[0].providerIdCriterion.providerId, PROVIDER)
+
+ def test_error_has_no_documents(self):
+ response = ingestion_pb2.QueryRequestStatusResponse()
+ response.exceptionalResult.message = "criteria list must not be empty"
+ self.stub.queryRequestStatus.return_value = response
+ result = self.client.query_request_status([RS.provider_id(PROVIDER)])
+ self.assertTrue(result.result_status.is_error)
+ self.assertIsNone(result.request_statuses)
+
+ def test_receive_limit_points_at_a_narrower_query(self):
+ self.stub.queryRequestStatus.side_effect = _FakeRpcError(
+ grpc.StatusCode.RESOURCE_EXHAUSTED, "Received message larger than max (5000000 vs. 4194304)"
+ )
+ message = self.client.query_request_status([RS.provider_id(PROVIDER)]).result_status.message
+ self.assertTrue(message.startswith("gRPC error: Received message larger than max"))
+ self.assertIn("narrow the query", message)
+
+
+class _FakeClock:
+ """Stands in for time.monotonic / time.sleep, so polling tests take no real time."""
+
+ def __init__(self):
+ self.now = 1_000.0
+ self.sleeps = []
+
+ def monotonic(self):
+ return self.now
+
+ def sleep(self, seconds):
+ self.sleeps.append(seconds)
+ self.now += seconds
+
+
+class TestAwaitRequestStatuses(_ClientTestCase):
+ def setUp(self):
+ super().setUp()
+ self.clock = _FakeClock()
+ patcher = patch.multiple(
+ "dp_python_lib.client.ingestion_client.time", monotonic=self.clock.monotonic, sleep=self.clock.sleep
+ )
+ patcher.start()
+ self.addCleanup(patcher.stop)
+
+ def _await(self, ids, **kwargs):
+ return self.client.await_request_statuses(PROVIDER, ids, since=SINCE, **kwargs)
+
+ def test_returns_each_ids_documents_after_one_poll(self):
+ self.stub.queryRequestStatus.return_value = _status_response(
+ _document("a"), _document("b", IngestionRequestStatus.ERROR), _document("unrelated")
+ )
+ statuses = self._await(["a", "b"])
+ self.assertEqual(list(statuses), ["a", "b"])
+ self.assertEqual(statuses["b"][0].ingestionRequestStatus, IngestionRequestStatus.ERROR)
+ self.assertEqual(self.clock.sleeps, [])
+
+ def test_a_repeated_id_returns_every_document(self):
+ self.stub.queryRequestStatus.return_value = _status_response(_document("a"), _document("a"))
+ self.assertEqual(len(self._await(["a"])["a"]), 2)
+
+ def test_one_query_per_poll_whatever_the_id_count(self):
+ ids = [f"chunk-{i}" for i in range(500)]
+ self.stub.queryRequestStatus.return_value = _status_response(*(_document(i) for i in ids))
+ self._await(ids)
+ self.assertEqual(self.stub.queryRequestStatus.call_count, 1)
+
+ def test_query_is_provider_and_a_skew_adjusted_floor(self):
+ self.stub.queryRequestStatus.return_value = _status_response(_document("a"))
+ self._await(["a"])
+ criteria = self.stub.queryRequestStatus.call_args[0][0].criteria
+ self.assertEqual(criteria[0].providerIdCriterion.providerId, PROVIDER)
+ time_range = criteria[1].timeRangeCriterion
+ expected_floor = (int(SINCE.timestamp()) - int(REQUEST_STATUS_CLOCK_SKEW.total_seconds())) * 1_000_000_000
+ self.assertEqual(to_epoch_nanos(time_range.beginTime), expected_floor)
+ # No end: the server's "now", rather than this machine's clock.
+ self.assertFalse(time_range.HasField("endTime"))
+ # No request-id criterion: ids are matched here, since the API takes one id per criterion.
+ self.assertFalse(any(c.HasField("requestIdCriterion") for c in criteria))
+
+ def test_floor_is_clamped_to_what_the_server_accepts(self):
+ self.stub.queryRequestStatus.return_value = _status_response(_document("a"))
+ self.client.await_request_statuses(PROVIDER, ["a"], since=10)
+ begin = self.stub.queryRequestStatus.call_args[0][0].criteria[1].timeRangeCriterion.beginTime
+ self.assertEqual((begin.epochSeconds, begin.nanoseconds), (1, 0))
+
+ def test_polls_until_every_id_appears(self):
+ self.stub.queryRequestStatus.side_effect = [
+ _status_response(),
+ _status_response(_document("a")),
+ _status_response(_document("a"), _document("b")),
+ ]
+ statuses = self._await(["a", "b"], poll_interval=0.5)
+ self.assertEqual(set(statuses), {"a", "b"})
+ self.assertEqual(self.clock.sleeps, [0.5, 0.5])
+
+ def test_timeout_names_the_missing_ids(self):
+ self.stub.queryRequestStatus.return_value = _status_response(_document("a"))
+ with self.assertRaises(TimeoutError) as ctx:
+ self._await(["a", "b", "c"], timeout=1.0, poll_interval=0.4)
+ message = str(ctx.exception)
+ self.assertIn("2 of 3", message)
+ self.assertIn("b, c", message)
+ self.assertNotIn("a,", message)
+ # The last sleep is trimmed to the deadline, then one final poll runs.
+ self.assertAlmostEqual(sum(self.clock.sleeps), 1.0)
+
+ def test_a_failed_query_raises(self):
+ self.stub.queryRequestStatus.side_effect = _FakeRpcError(grpc.StatusCode.UNAVAILABLE, "down")
+ with self.assertRaises(RuntimeError) as ctx:
+ self._await(["a"])
+ self.assertIn("gRPC error: down", str(ctx.exception))
+
+ def test_duplicate_input_ids_are_waited_for_once(self):
+ self.stub.queryRequestStatus.return_value = _status_response(_document("a"))
+ self.assertEqual(list(self._await(["a", "a"])), ["a"])
+
+ def test_argument_validation(self):
+ for label, call in (
+ ("blank provider", lambda: self.client.await_request_statuses(" ", ["a"], since=SINCE)),
+ ("no ids", lambda: self._await([])),
+ ("blank id", lambda: self._await(["a", ""])),
+ ("zero timeout", lambda: self._await(["a"], timeout=0)),
+ ("zero interval", lambda: self._await(["a"], poll_interval=0)),
+ ):
+ with self.subTest(case=label), self.assertRaises(ValueError):
+ call()
+
+ def test_since_is_required(self):
+ with self.assertRaises(TypeError):
+ self.client.await_request_statuses(PROVIDER, ["a"]) # type: ignore[call-arg]
+
+
if __name__ == "__main__":
unittest.main()
diff --git a/tests/unit/test_ingestion_streaming_grpc.py b/tests/unit/test_ingestion_streaming_grpc.py
new file mode 100644
index 0000000..7c27971
--- /dev/null
+++ b/tests/unit/test_ingestion_streaming_grpc.py
@@ -0,0 +1,217 @@
+"""
+The streaming ingest methods against a real, in-process grpcio server (#17, D5).
+
+A mocked stub cannot show the behavior these tests exist for: grpcio consumes a request iterator on its own thread,
+and when the iterator raises it cancels the call and reports only UNKNOWN "Exception iterating requests!". The
+client must re-raise the caller's own exception instead. The server here listens on an ephemeral localhost port
+and records what it received, so the tests can also check that requests sent before the failure did arrive.
+"""
+
+import contextlib
+import os
+import sys
+import threading
+import time
+import unittest
+from concurrent import futures
+
+import grpc
+
+# Add src directory to path for imports
+sys.path.insert(0, os.path.join(os.path.dirname(__file__), "../../src"))
+
+from dp_python_lib.client import data_frame as dfb
+from dp_python_lib.client.ingestion_client import IngestDataRequestParams, IngestionClient
+from dp_python_lib.grpc import ingestion_pb2, ingestion_pb2_grpc
+
+REJECTED_ID = "reject-me"
+
+
+class _RecordingServicer(ingestion_pb2_grpc.DpIngestionServiceServicer):
+ """Acks every request except REJECTED_ID, and records the ids it received."""
+
+ def __init__(self):
+ self.received: list[str] = []
+ self._lock = threading.Lock()
+
+ def _record(self, request):
+ with self._lock:
+ self.received.append(request.clientRequestId)
+
+ @staticmethod
+ def _response(request):
+ response = ingestion_pb2.IngestDataResponse(
+ providerId=request.providerId, clientRequestId=request.clientRequestId
+ )
+ if request.clientRequestId == REJECTED_ID:
+ response.exceptionalResult.message = "rejected by test server"
+ else:
+ response.ackResult.numRows = request.ingestionDataFrame.dataTimestamps.samplingClock.count
+ response.ackResult.numColumns = len(request.ingestionDataFrame.doubleColumns)
+ return response
+
+ def ingestDataStream(self, request_iterator, context):
+ ids, rejected = [], []
+ for request in request_iterator:
+ self._record(request)
+ ids.append(request.clientRequestId)
+ if request.clientRequestId == REJECTED_ID:
+ rejected.append(request.clientRequestId)
+ response = ingestion_pb2.IngestDataStreamResponse(clientRequestIds=ids, rejectedRequestIds=rejected)
+ if rejected:
+ response.exceptionalResult.message = "one or more requests were rejected"
+ else:
+ response.ingestDataStreamResult.numRequests = len(ids)
+ return response
+
+ def ingestDataBidiStream(self, request_iterator, context):
+ for request in request_iterator:
+ self._record(request)
+ yield self._response(request)
+
+
+def _params(request_id):
+ frame = dfb.data_frame(dfb.sampling_clock(1_770_000_000, 1_000, 2), [dfb.double_column("pv", [1.0, 2.0])])
+ return IngestDataRequestParams("provider-1", frame, client_request_id=request_id)
+
+
+class _Boom(Exception):
+ """The caller's own exception type, which the client must hand back unchanged."""
+
+
+def _failing_after(count):
+ """A request generator that yields `count` requests and then raises _Boom."""
+ for index in range(count):
+ yield _params(f"ok-{index}")
+ raise _Boom("producer failed")
+
+
+class TestStreamingAgainstInProcessServer(unittest.TestCase):
+ def setUp(self):
+ self.servicer = _RecordingServicer()
+ self.server = grpc.server(futures.ThreadPoolExecutor(max_workers=4))
+ ingestion_pb2_grpc.add_DpIngestionServiceServicer_to_server(self.servicer, self.server)
+ port = self.server.add_insecure_port("localhost:0")
+ self.server.start()
+ self.channel = grpc.insecure_channel(f"localhost:{port}")
+ self.client = IngestionClient(self.channel)
+
+ def tearDown(self):
+ self.channel.close()
+ self.server.stop(grace=None)
+
+ # ---- the grpcio behavior D5 works around
+
+ def test_grpcio_itself_hides_the_iterator_exception(self):
+ # Pins the premise: through the raw stub, the caller's exception is gone. If grpcio ever starts chaining or
+ # re-raising it, this fails, and _RequestFeed can be reconsidered.
+ stub = ingestion_pb2_grpc.DpIngestionServiceStub(self.channel)
+
+ def requests():
+ yield ingestion_pb2.IngestDataRequest(providerId="p", clientRequestId="r")
+ raise _Boom("producer failed")
+
+ with self.assertRaises(grpc.RpcError) as ctx:
+ stub.ingestDataStream(requests())
+ self.assertEqual(ctx.exception.code(), grpc.StatusCode.UNKNOWN)
+ self.assertNotIsInstance(ctx.exception.__cause__, _Boom)
+
+ # ---- ingestDataStream
+
+ def test_stream_happy_path(self):
+ result = self.client.ingest_data_stream(_params(f"id-{i}") for i in range(3))
+ self.assertFalse(result.result_status.is_error)
+ self.assertEqual(result.num_requests, 3)
+ self.assertEqual(result.client_request_ids, ["id-0", "id-1", "id-2"])
+ self.assertEqual(result.rejected_request_ids, [])
+
+ def test_stream_partial_reject_keeps_the_response(self):
+ result = self.client.ingest_data_stream([_params("a"), _params(REJECTED_ID), _params("b")])
+ self.assertTrue(result.result_status.is_error)
+ self.assertEqual(result.rejected_request_ids, [REJECTED_ID])
+ self.assertEqual(result.client_request_ids, ["a", REJECTED_ID, "b"])
+ self.assertIsNone(result.num_requests)
+
+ def test_stream_reraises_the_callers_exception(self):
+ with self.assertRaises(_Boom) as ctx:
+ self.client.ingest_data_stream(_failing_after(2))
+ self.assertEqual(str(ctx.exception), "producer failed")
+ self.assertIsInstance(ctx.exception.__cause__, grpc.RpcError)
+ if sys.version_info >= (3, 11):
+ self.assertTrue(any("2 request(s) had been handed to gRPC" in note for note in ctx.exception.__notes__))
+ # Handed over is not delivered: grpcio's cancellation races the sends, so the server holds some prefix of
+ # the requests produced before the failure -- which is why the note says "any that reached the server".
+ self.assertIn(self.servicer.received, (["ok-0", "ok-1"], ["ok-0"], []))
+
+ def test_stream_reraises_when_the_first_request_fails(self):
+ with self.assertRaises(_Boom):
+ self.client.ingest_data_stream(_failing_after(0))
+ self.assertEqual(self.servicer.received, [])
+
+ # ---- ingestDataBidiStream
+
+ def test_bidi_yields_one_result_per_request_including_rejects(self):
+ results = list(self.client.iter_ingest_data_bidi_stream([_params("a"), _params(REJECTED_ID), _params("b")]))
+ self.assertEqual([r.client_request_id for r in results], ["a", REJECTED_ID, "b"])
+ self.assertEqual([r.result_status.is_error for r in results], [False, True, False])
+ self.assertEqual(results[1].result_status.message, "rejected by test server")
+ self.assertEqual((results[0].num_rows, results[0].num_columns), (2, 1))
+
+ def test_bidi_reraises_the_callers_exception(self):
+ seen = []
+ with self.assertRaises(_Boom) as ctx:
+ for result in self.client.iter_ingest_data_bidi_stream(_failing_after(2)):
+ seen.append(result.client_request_id)
+ self.assertIsInstance(ctx.exception.__cause__, grpc.RpcError)
+ # As for the client stream, delivery of the requests before the failure races the cancellation, and so
+ # does the arrival of their acks.
+ self.assertIn(self.servicer.received, (["ok-0", "ok-1"], ["ok-0"], []))
+ self.assertIn(seen, (["ok-0", "ok-1"], ["ok-0"], []))
+
+ def _gated_producer(self):
+ """1,000 requests whose producer blocks after three until the gate opens, and a record of what it made."""
+ gate = threading.Event()
+ produced = []
+
+ def gated():
+ for index in range(1_000):
+ if index == 3:
+ gate.wait(5)
+ produced.append(index)
+ yield _params(f"n-{index}")
+
+ return gated(), gate, produced
+
+ def _assert_stopped_sending(self, gate, produced):
+ gate.set()
+ time.sleep(0.3) # time enough for an uncancelled call to drain all 1,000
+ self.assertLess(len(produced), 10)
+ self.assertLess(len(self.servicer.received), 10)
+
+ def test_bidi_stops_sending_when_the_results_are_closed(self):
+ # grpcio pulls requests on its own thread. The producer blocks after a few requests until the caller has
+ # closed the results; without a cancel, grpcio would then drain the rest of it into the server.
+ requests, gate, produced = self._gated_producer()
+ results = self.client.iter_ingest_data_bidi_stream(requests)
+ next(results)
+ results.close()
+ self._assert_stopped_sending(gate, produced)
+
+ def test_bidi_stops_sending_on_a_break_inside_closing(self):
+ # The documented way to stop early. A bare break would not do: the loop's iterator stays referenced, so it
+ # is not closed until collected, and grpcio keeps sending meanwhile.
+ requests, gate, produced = self._gated_producer()
+ with contextlib.closing(self.client.iter_ingest_data_bidi_stream(requests)) as results:
+ for _result in results:
+ break
+ self._assert_stopped_sending(gate, produced)
+
+ def test_bidi_transport_failure_raises_runtime_error(self):
+ self.server.stop(grace=None)
+ with self.assertRaises(RuntimeError) as ctx:
+ list(self.client.iter_ingest_data_bidi_stream([_params("a")]))
+ self.assertIn("gRPC error", str(ctx.exception))
+
+
+if __name__ == "__main__":
+ unittest.main()
diff --git a/tests/unit/test_service_api_client_base.py b/tests/unit/test_service_api_client_base.py
index 6d2f1d9..7f28e62 100644
--- a/tests/unit/test_service_api_client_base.py
+++ b/tests/unit/test_service_api_client_base.py
@@ -172,6 +172,27 @@ def test_default_logs_used_when_callables_omitted(self):
self.assertIn("Calling someOperation API", "\n".join(captured.output))
self.assertIn("someOperation completed successfully", "\n".join(captured.output))
+ def test_rpc_error_hint_is_appended_to_the_grpc_error_message(self):
+ error = grpc.RpcError()
+ error.details = Mock(return_value="too big")
+ stub_call = Mock(side_effect=error)
+ hint = Mock(return_value=" (split it)")
+
+ result = self._dispatch(stub_call, rpc_error_hint=hint)
+
+ # Appended, never substituted: the "gRPC error: " prefix is part of the result contract.
+ self.assertEqual("gRPC error: too big (split it)", result.result_status.message)
+ hint.assert_called_once_with(error)
+
+ def test_rpc_error_hint_not_consulted_on_success_or_business_error(self):
+ hint = Mock(return_value=" (unused)")
+ for field in ("someResult", "exceptionalResult"):
+ with self.subTest(field=field):
+ response = _response_with_field(field)
+ response.exceptionalResult.message = "rejected"
+ self._dispatch(Mock(return_value=response), rpc_error_hint=hint)
+ hint.assert_not_called()
+
if __name__ == "__main__":
unittest.main()
diff --git a/tests/unit/test_time_conversions.py b/tests/unit/test_time_conversions.py
index ef89788..6b5665f 100644
--- a/tests/unit/test_time_conversions.py
+++ b/tests/unit/test_time_conversions.py
@@ -6,7 +6,7 @@
# Add src directory to path for imports
sys.path.insert(0, os.path.join(os.path.dirname(__file__), "../../src"))
-from dp_python_lib.client.time_conversions import to_epoch_nanos, to_timestamp
+from dp_python_lib.client.time_conversions import from_epoch_nanos, to_epoch_nanos, to_timestamp
from dp_python_lib.grpc import common_pb2
@@ -31,6 +31,23 @@ def test_zero_timestamp_is_zero(self):
self.assertEqual(to_epoch_nanos(common_pb2.Timestamp()), 0)
+class TestFromEpochNanos(unittest.TestCase):
+ """from_epoch_nanos() is to_epoch_nanos()'s inverse, shared by three modules since split_data_frame() (#17)."""
+
+ def test_splits_into_seconds_and_nanoseconds(self):
+ ts = from_epoch_nanos(1_770_055_200_123_456_789)
+ self.assertEqual((ts.epochSeconds, ts.nanoseconds), (1_770_055_200, 123_456_789))
+
+ def test_round_trips_with_to_epoch_nanos_exactly(self):
+ for original in (0, 1, 999_999_999, 1_000_000_000, 1_770_055_200_123_456_789):
+ with self.subTest(original=original):
+ self.assertEqual(to_epoch_nanos(from_epoch_nanos(original)), original)
+
+ def test_negative_is_rejected_by_the_unsigned_field(self):
+ with self.assertRaises(ValueError):
+ from_epoch_nanos(-1)
+
+
class TestDatetimeConversionIsExact(unittest.TestCase):
"""
A datetime must not route through datetime.timestamp().