Skip to content

Add reusable prepared datasets - #34

Merged
samueleternity merged 11 commits into
mainfrom
dataset
Sep 27, 2026
Merged

samueleternity merged 11 commits into
mainfrom
dataset

Conversation

@samueleternity

Copy link
Copy Markdown
Owner

This adds save/load support for prepared datasets so training and inference can reuse processed data without re-tokenizing raw sources. It introduces --save, --load-prepared, and --prepare-only flows, persists the resolved dataset state, and wires prepared datasets through task construction while validating type compatibility and path safety.

This adds save/load support for prepared datasets so training and inference can reuse processed data without re-tokenizing raw sources. It introduces --save, --load-prepared, and --prepare-only flows, persists the resolved dataset state, and wires prepared datasets through task construction while validating type compatibility and path safety.
Introduce an additive classic-task training track for long-window prediction and recall probes across text/audio/video and multimodal sources. This adds the dataset family, CLI flags, registry support, OGS/KL attribution reporting, classic inference generation, and compatibility for evaluating memory-on/off probe behavior while preserving the existing KV-chain task flow.
Add the src root to sys.path so the training module works when launched either as `python -m initium.core_training` or directly from a source checkout. This keeps imports consistent regardless of execution context.
Classic datasets now accept runtime overrides for window size, probe distances, gamma, and OOD eval episodes even when loaded from prepared data. Training also adds a --batch-size option with a classic default of 1 to keep full-window DNC activations within memory, and the generated run ID reflects the selected classic batch size. The classic inference path applies the same runtime settings to prepared streams, and the output projection path no longer requires a contiguous tensor.
This change avoids expensive per-step host/device syncs during training by keeping scalar loss terms on-device until logging intervals and computing averaged values from accumulated tensors. Finite-value checks are now limited to the logging cadence, reducing noisy overhead in classic-track runs.

It also tightens the classic task metrics path by keeping loss terms as tensors and computing diversity without per-row Python object conversion, and updates the NVRTC compatibility shim to avoid repeated CUDA-to-CPU syncs on non-negative DNC inputs while adding a strict debug mode for negative-input detection.
Classic datasets now encode each token with the existing 3-digit codec instead of a 1000-wide one-hot slot, while preserving a separate per-modality probe marker. This updates dataset sizing, window construction, generated-token encoding, and the classic track diagnostic output, and clarifies the CUDA no-sync compatibility patch message.
Add DigitCodec.decode_digits to convert a sequence of digit IDs (tensor or list) into an integer label with input validation. Replace calls to decode_field on target digit arrays in ClassicDataset and ClassicInferenceTask with decode_digits so already-decoded ID sequences are handled correctly (avoids treating them as logits/floats). Improves error messages for wrong lengths or out-of-range digit IDs.
This change turns the package root into a lazy, stable public API for architecture components and controller wrappers. It centralizes exports via __getattr__ so optional training and inference dependencies are not imported eagerly, keeps internal data encoders and task implementations out of the public contract, and adds checkpoint/model loading helpers for easier integration.
This removes the additive classic long-window dataset/task family and its CLI flags, including the classic configs, dataset registry entries, inference generation paths, and training logs tied to memory probe evaluation. It also simplifies OGS scoring to the core offset/memory/perfect components and updates attribution reporting to be architecture-agnostic rather than classic-specific.
Add an explicit type annotation for final_ood_field_log in src/initium/core_training.py: dict[tuple[int,int], list[int]]. This improves type clarity and helps static type checkers; no functional behavior changed.
@samueleternity
samueleternity merged commit 29465e9 into main Sep 27, 2026
5 checks passed
Sign up for free to join this conversation on GitHub. Already have an account? Sign in to comment

Labels

None yet

Projects

None yet

Development

Successfully merging this pull request may close these issues.

add more usable training mechanism add more optionality for dataset usage

1 participant