Repository navigation
RDF 1.2 triple terms: rdf:reifies links, triple terms as values, faster annotation queries - #2017
Conversation
OPST holds every object type, segmented by o_type. IndexType::for_query gates it on o_is_ref only as a conservative default, and the binary scan path already takes OPST for any constant object.
A triple term `<<( s p o )>>` becomes a first-class object value. `FlakeValue::TripleTerm` carries the materialized form; the index stores a dictionary handle under `OType::TRIPLE_TERM` (tag-10, decode kind `TripleTermDict`, V1 kind byte 0x15). Handles are partitioned by the inner predicate (`triple_term::term_handle`) so a predicate restriction is one `o_key` interval, and `TermKey` is the encoded base edge, ordered subject-first for the reverse tree. Adds the `rdf:reifies` and `f:tripleTerm` sids, plus a read-side `is_scan_hidden_predicate` so wildcard scans keep hiding the link form alongside the `f:reifies*` bundle while it is index-internal. Every closed match over `FlakeValue` gains an arm: formatters render the term by its display form for now, and the commit codec rejects the value until the write path moves over.
The link `_:r rdf:reifies <<( s p o )>>` is now one main-index flake on imported ledgers, alongside the `f:reifies*` bundle it is meant to replace. Its object is a handle into a ledger-global term dictionary that mirrors the subject dictionary: one forward pack stream per inner predicate keyed by sequence, one reverse tree keyed by the encoded base edge, and references in a flagged trailing root section (`FLAG_EXT_HAS_TERM_DICT`) that the core metadata parser skips. Import interns terms after the global remap, where ids are final. The sink records each reified edge as a pseudo-record in a per-chunk term table and spools the link with the entry's ordinal; both build phases resolve the ordinal through `TermRemapCtx`, which remaps the entry with the chunk's own tables and interns it in a build-wide `TermDictBuilder`. The dictionary is uploaded next to the index artifacts and referenced from the root, and the store decodes a handle back to the materialized term. Full rebuilds and incremental builds do not intern yet (a rebuild drops the term section), and the query lowering still takes the bundle path.
…builds The resolver now assembles each commit's `f:reifies*` bundle into the RDF 1.2 link record and a per-chunk term entry (`link_synth`), so an index built from commits carries the same `rdf:reifies` flakes bulk import writes. This is what makes the CLI's post-import seal, and every later reindex, keep the links and the dictionary. Full rebuilds intern in Phase C once ids are global and publish the dictionary from the root. Incremental builds look each term up in the base dictionary, allocate fresh handles above its per-predicate watermarks, append per-predicate forward packs and update the reverse tree copy-on-write, with the replaced tree blobs recorded as garbage. Pack appends take the pack kind explicitly, since the subject appender hardcoded it. The stage cascade drops the index-side link before partitioning a reifier's flakes into bundle and body, so an indexed reifier whose body is deleted still looks orphaned and cascades; the link itself is retired by the next index pass. Tests cover import, reindex and the incremental window, including a second reifier on an existing edge.
Behind `FLUREE_ANNOTATION_TERMS=1`, the SPARQL lowering routes the
reified-triple forms (`<< s p o >>` and `?r rdf:reifies <<( s p o )>>`)
through the link flake instead of the `f:reifies*` bundle chain: one
`annotation rdf:reifies ?__term_N` triple, then `BIND(SUBJECT(?__term_N)
AS ?s)` for a component variable no earlier pattern binds, or a `FILTER`
equality when one does, since `BIND` overwrites. A constant predicate
becomes `FILTER(PREDICATE(?__term_N) = <p>)`, which the filter pushdown
turns into `ObjectBounds::term_predicate` and the scan into one `POST`
handle interval: the seek keys narrow the leaflets, and each row's
`o_key` is checked against the interval in encoded form, so a count
never decodes a term. A fully constant edge composes to a constant term
the scan resolves through the reverse tree.
`SUBJECT`, `PREDICATE`, `OBJECT` and `isTRIPLE` now evaluate over
`FlakeValue::TripleTerm`; `TRIPLE()` stays deferred. Link scans emit
late-materialized handles. The annotation syntax `s p o {| |}` and the
JSON-LD surface still take the bundle chain.
The link tests move to their own binary because the flag is
process-wide; the new test pins every access-path shape from the design
doc, including a component variable bound by an earlier pattern.
… the link lowering
The flagged link lowering was slower than the bundle chain: a composed
constant term never reached the dictionary (the simple encoder had no
TripleTerm arm, so a fully bound pattern fell to a decode-and-compare
scan of every link), every SUBJECT/PREDICATE/OBJECT accessor materialized
the whole term per row, and positions nothing reads were still bound.
- compose_term_handle is shared by both encoders; an uninterned term is
NotFound, so the scan skips straight to novelty.
- An accessor on a late-materialized handle reads the dictionary key and
binds the component encoded (EncodedSid, EncodedPid, EncodedLit with
its datatype and tag); OBJECT on a materialized term keeps the
datatype and tag too.
- elide_unread_term_binds drops component BINDs nothing reads, the twin
of elide_redundant_chain for the bundle chain.
- Variable positions are always BIND: the agreement check is the join in
every scope, and bind_unifies normalizes representations (a scan's
EncodedSid against a VALUES row's Sid) in BindOperator and apply_inline.
The lowering's cross-scope binding set is gone.
- Constant positions are sameTerm filters the pushdown turns into
ObjectBounds::{term_subject, term_object} beside term_predicate; the
scan checks them on the handle's dictionary key (term_key_filter). A
literal object carries its datatype or tag as a resolved binding, so
"5"^^xsd:int and "chat"@fr match by term identity, as the composed
constant does. Two constraints on one component that disagree keep the
second filter instead of one silently winning.
- Import skips arena-scoped (decimal, big-integer, vector) objects as
rebuild does: per-(graph, predicate) handles are no graph-independent
term identity.
…undles per commit A re-point that changes one slot writes only that slot's retract and assert (sync and upsert cancel the unchanged slots), and a full re-point's six ops interleave under the per-commit sort, so an assembler that needed all three slots in one (graph, reifier, t, op) group emitted no link events and the old link survived the build. The resolver now only collects f:reifies* ops (RebuildChunk.attachments) and the build replays them per reifier once ids are global: ops in t order, retracts before asserts within one t, a link retract and assert wherever the attachment changes from or to a complete edge. A rebuild replays the whole history from nothing and writes the links as one more sorted commit file; an incremental build seeds each touched reifier with the attachment the base index holds (one SPOT point lookup) and appends the links to the window.
…s look The count planner and the other fused fast paths classify the query's pattern list as lowered, so a `COUNT(*)` over `<< ?s ?p ?o >>` still carried three component BINDs at that point and fell through to the generic join, ten times slower than the same two-triple join written by hand. Run the unread-position elision on the query before any fast path sees it; the WHERE planner's own pass stays for the seeded and nested cases.
OBJECT() of a materialized term, and DATATYPE/LANG over any accessor, went through the numeric comparable, which carries no datatype: an indexed "5"^^xsd:int read as xsd:integer as soon as an unrelated unindexed commit turned late materialization off, and DATATYPE(OBJECT(?t)) lost the subtype even on a quiescent ledger. term_component_binding now yields the component as a binding on both the encoded path (dictionary key) and the materialized path (the term itself), with datatype and tag, and DATATYPE and LANG read an accessor argument through it exactly as they read a bound variable. Pinned on both paths.
…nk replay The incremental replay read a reifier's base predicate slot with a swallowed error, so a failed dictionary fetch looked like a missing predicate, the replay saw an incomplete attachment before and after an object-only re-point, and the old link survived. The read error and a predicate absent from the dictionary now fail the build. Rebuild frees each chunk's record and term buffers as soon as its spool file is written instead of after the loop, and the replay hands link records to the spool one reifier at a time, in (graph, SPOT) order, rather than accumulating them beside the attachment table.
A reified-triple pattern lowers to `?r rdf:reifies ?t`, and a row such as `ex:r rdf:reifies ex:x` satisfied it: the accessor BINDs failed and kept the row, and a count elided them altogether. Neither a filter nor a datatype constraint on the link triple is an option, since both take the shape off the count plan and the planner paths that gate on dtc.is_none(), and the count planner counts from leaf-directory metadata that cannot see object types. So the predicate carries the rule. The write path refuses rdf:reifies with a non-term object on every surface (JSON-LD, SPARQL UPDATE, Turtle transact, the template path, bulk import), as it refuses user-authored f:reifies*; the reified-triple forms never reach those checks. The scan treats an rdf:reifies pattern with a variable object as term-only on the encoded object type, on novelty flakes too, and narrows a POST seek to the term segment. Rows written as a plain predicate before this rule are dropped by the scan and still counted by the count plan.
An incremental build replays each touched reifier from the attachment the base index holds, and a re-point retracts the old term's link. Both lookups ran once per reifier: a fresh SPOT cursor for the base attachment, and a reverse-tree lookup for each term, which rereads a whole ~2.3 MB leaf when no leaf cache is attached. The base attachments now come from one batched SPOT pass per graph (batched_lookup_subject_properties), with each predicate IRI resolved once. Term handles are resolved together: a dry replay names the terms the real replay will ask for, and TermDictReader::find_handles reads each leaf once. On the 10% annotation benchmark slice, an incremental index over 100k partial re-points drops from 72.7 s to 30.2 s (the bundle path on main: 27.0 s), at the same peak memory.
`fluree_db_core::io_stats` counts artifact reads by kind and source, and the CLI prints the table on exit when FLUREE_IO_STATS is set. Remote fetches are counted once, at the storage bridge (`store` / `store-range`), with dictionary blobs told apart by format (string, subject and term packs, reverse-tree leaves and branches); readers count what a local path or the disk artifact cache served. FLUREE_FORCE_REMOTE_READS makes a local store answer as a remote one: no resolved local paths and no resident bytes, so every artifact goes through `get` / `get_range` and the disk cache, and the counts are what S3 would be asked for. Off, each hook is one relaxed atomic load and builds no label.
OType::TRIPLE_TERM joined the Fluree-reserved entries, so the built-in table holds 45 entries, not 44.
…al builds Phase 3b seeds every class entry of the base stats and resolves each class Sid to its sid64 through the subject reverse tree, one lookup per class. Without a leaf cache each lookup rereads a whole ~2.3 MB leaf, so a ledger with many classes paid it on every incremental pass: 203,269 classes took 21 s on the 10% annotation benchmark slice, whatever the size of the window. `BinaryIndexStore::find_subject_ids_by_parts` resolves them together through `reverse_lookup_many`, reading each leaf once. A failed read now aborts to a full rebuild, as the base PSOT class lookup beside it does, instead of dropping the class from the published stats.
Loading an indexed ledger read the root twice: `LedgerState::load` for the snapshot metadata, then the binary store load for the full decode. Against a remote store that is a second whole-object GET of the largest single artifact on every cold load (83 MB twice on the annotation benchmark slice). `LedgerState::load_with_root` / `load_with_store_and_root` hand back the root they read, and `load_and_attach_binary_store` takes it when its CID is the one the record names. Every load-then-attach path passes it through.
The root's stats carry per-graph class tables and a ledger-wide one, and both build paths set the ledger-wide table to `union_per_graph_classes` of the per-graph ones. On a ledger with many classes the two copies are most of the root: 41.5 MB each for 203,269 classes on the annotation benchmark slice. The encoder now writes the ledger-wide table only when no graph carries classes, and both decoders derive it from the graphs when it is absent. Roots that still carry it decode as before. Import roots, which never set it, now present the same class table a rebuild would, instead of none until the next reindex. A reader without this change sees no ledger-wide class table on a new root and plans without class statistics.
Each class-property usage spelled out its property Sid, and each ref-class count its class Sid, with fixed-width counts. On a ledger with many classes that is nearly the whole root: 203,269 classes, 758,489 usages and 347,070 ref counts took 41.5 MB on the annotation benchmark slice, though only 607 predicates appear among the usages. The class tables now travel in an appended stats section (tag 2): every Sid written once in a sorted, front-coded table, entries referring to it by index, counts as varints. Re-encoding the slice's 83 MB root gives 8.4 MB with identical decoded stats, and decoded entries share their Sid names. Readers skip appended sections they do not know, so the legacy per-graph class slots are left empty and a reader without this change sees no class tables rather than misparsing. Roots that carry classes in the legacy slots still decode. The binary-index stats decoder now delegates to the core one, so the format has a single parser.
The state load decoded the root for the snapshot metadata (`LedgerSnapshot::from_root_bytes`) and the binary store load decoded the same bytes again as an `IndexRoot`; on the annotation benchmark slice each decode took about 110 ms. `LedgerState::load_decoding_root` / `load_with_store_decoding_root` take the root decoder from the caller and hand back what it kept. The API decodes the root once as an `IndexRoot`, builds the snapshot from it through `IndexRoot::take_snapshot_metadata` (stats and schema move into the snapshot rather than being copied), and passes the rest of the root to the binary store load. Index publish builds its snapshot through the same method. `load` and `load_with_store` keep the metadata-only decode.
Each incremental build appended a pack to every inner predicate's term stream that gained terms, and nothing merged them, so a predicate's routing table grew by one entry per build forever. The streams now go through the same tail-first compaction as the subject and string streams, within the cycle's shared budget (hoisted so every stream draws on one allowance). Packs a merge consumes, base packs and this cycle's own uploads alike, are recorded as garbage with the term dictionary's replaced reverse-tree leaves.
…aries The server's background warmer touched string and subject packs only, so the first query to decode reified edges fetched every term pack it needed one at a time. Term packs are now warmed last, within the same budget.
A BIND into a variable that an earlier step already bound is an equality check against that column. When the bind's expression waits on a later pattern, the planner holds it back, and each intermediate join step trims its output to the variables still live: later triples, pending filters, and the variables pending binds' expressions read. The bind's own target was not among them, so a target nothing else read was trimmed, the bind found no value and bound afresh, and the check disappeared: a count became the cross product. JSON-LD reaches it with a rebind (`["bind", "?a", "?v"]` after `?a` is bound and one more join): 6 rows instead of 1. The reified-edge lowering reaches it from SPARQL, since every quoted position is a BIND and a shared component variable is joined that way: `<< ex:a ex:p ?o1 >> ... << ?o1 ex:q ?o2 >> ...` under COUNT(*) counted 167,551,488 instead of 189 on the annotation benchmark slice. A pending bind now keeps its target live.
…ing templates
The write firewall checked `rdf:reifies` where a surface spells the
predicate out. A predicate variable is resolved only when a template is
materialized per solution row, so `INSERT { ex:r ?p ex:x } WHERE { VALUES
?p { rdf:reifies } }`, or the JSON-LD update with a `?p` key, wrote a
non-term link row. The count plan answers from leaf metadata, which would
count it, while the scan drops it.
Template materialization now refuses an asserted `rdf:reifies` flake whose
resolved object is not a triple term, which covers SPARQL UPDATE, JSON-LD
update and every other template path. Retracting a legacy row still works.
`term_component_binding` produced a binding, with the object's datatype or tag, only when the accessor's argument was a bare variable. Any other argument (`OBJECT(COALESCE(?t))`) fell through to the numeric comparable, which carries no datatype, so `DATATYPE` reported `xsd:integer` for an `xsd:int` object and a `BIND` over the accessor lost the subtype too. The general path now evaluates the argument to a term and builds the same binding; the encoded-handle shortcut for a bare variable is unchanged, and the comparable accessor converts that binding like any other.
…weep Term streams were compacted only in a cycle that added terms to them, after the string and subject streams had drawn on the shared budget, so a term stream that missed its turn and then went quiet stayed fragmented for good. The term dictionary now reserves a third of each cycle's compaction budget before the other dictionaries run, and takes over what they leave. Its streams that grew are compacted first, then the rest in a rotation keyed by the base index t, as the subject maintenance sweep does; a cycle with no new terms still runs the sweep and publishes the merged routing table.
…later one `DictTreeReader::reverse_range_scan` found its first leaf with `find_leaf`, which answers only for a key inside some leaf's [first, last] range. A start key in the gap between two leaves (past one leaf's last key, before the next one's first) found no leaf, and the scan returned nothing although the later leaves held matching keys. A subject prefix lookup (`find_subjects_by_prefix`) whose prefix sorts into such a gap missed every match. The scan now starts at the first leaf whose last key is not below the start key.
…tern The link lowering decomposed `<< s p o >>` into `?r rdf:reifies ?t` plus a BIND per variable position. Component equality was then visible only after the link scan: `<< X :treats ?o1 >> ... << ?o1 :causes ?o2 >> ...` read every `:causes` link once per left row and joined on `?o1` in the BIND check, 12 s against 90 ms for the bundle chain on the annotation benchmark slice. Quoted patterns now lower to the link plus `TermComponents(?t, s, p, o)`, an internal pattern the planner places like any other source. Its operator finds candidate terms per input row: - the term variable bound: one dictionary key decode supplies every component; - the subject bound, as a constant or by an earlier pattern: the reverse tree's subject-first `(s_id[, p_id])` prefix range, which yields keys and handles together; distinct prefixes are read once per batch; - neither: the terms of the bound predicate, or of the dictionary. A variable component another pattern bound first is joined on with the same equality a BIND check applied. The planner prices a bound term as a decode and an anchored subject as a small source, so the chained edge is found through its subject and its link becomes a bound-object lookup; the link keeps the graph, time, policy and novelty checks. Constant components stay filters on the link as well, so a link-first plan keeps its handle interval. Unread positions are elided as before, and a pattern left with nothing to bind or anchor goes, which keeps the count plan. On the slice the populated chained query takes 60-75 ms (bundle chain 92 ms), with every other measured shape unchanged or faster.
… covered Links existed only in the index, so a link query missed every annotation added, retracted or re-pointed since the last build. Novelty now derives them from each commit's f:reifies* slot ops, replaying a reifier from the index's attachment plus its earlier novelty ops, the rule the index build applies. The index's attachments come from a base the ledger installs: an empty one when the index never held an annotation, the binary store's otherwise (at load, at each publish, for historical views). Commits applied before the base is known wait and are derived when it arrives; a trim past the base drops it. Terms only novelty names take the raw-flake lane, and TermComponents offers them beside the dictionary's, also on a ledger with no index. Graph management and the lifecycle probe skip derived links: no transaction writes one.
A link whose term the index has not interned took the raw-flake lane, so any annotation committed since the last build took a link count off the count plan (P2 on the dev slice: 25 ms indexed, 2.8 s with 10k such links). Dictionary novelty now keeps a term table, populated from the links novelty derives, and a term there gets a provisional handle in its inner predicate's interval, above every indexed sequence: 2^31 plus its index. The term-dictionary builder stops below that. Composition tries the persisted dictionary first. A provisional handle's key is computed from its novelty term, the novelty-aware graph view and the batched join decode it, and publish retires the entries the index now interns. apply_commit returns the links it derived so each write path registers them. P2 with the same novelty: 69 ms, on the count plan.
A scan whose bound object was a term walked every novelty flake of its predicate, so each chained link lookup re-translated all of novelty's links, composing each term (P19b over 10k unindexed annotations: 12 s). Terms order by base subject and predicate first, so the terms sharing the bound term's form one run, which the walk now brackets.
A triple-term object is checked as the triple it names, but the scan and probe lanes skip per-flake filtering whenever the scanned predicate is uncovered by the view policy. On an indexed ledger a policy hiding ex:salary therefore still returned the hidden triple through `?r rdf:reifies ?t`, `<< ?s ex:salary ?o >> ...` and any other predicate holding a term of it. Under a view policy that restricts anything, the shared predicate gate now declines `rdf:reifies` and any predicate whose observed datatypes include the untagged kind triple terms carry, are unknown, or miss novelty the stats cannot see. The policy test now runs from novelty, from an index and from novelty over one, and covers a term held under another predicate.
On an annotated ledger indexed by an earlier release, link queries refused with the reindex error, but export returned Ok with every reifier body orphaned and no markers, and a crawl dropped @annotation: both read links the index does not have. The snapshot now carries `needs_link_reindex` (annotations, no term dictionary), which also gives embedding hosts a machine-readable signal to schedule the rebuild. Queries, export and hydration's annotation reads all refuse through `require_link_index`, whose message now names the API route as well as the CLI one. A ledger written and indexed by 4.2.3 is checked in as a fixture to pin it.
…er side Past 1,024 distinct outer keys the semijoin fell back to an unseeded build over the whole body: an annotated edge's EXISTS on a large predicate cost the predicate, not the query (IC7, and 1,100 hub likes over 1M ex:likes at 2-3x main's time and ~11x its memory). And because the outer side was pulled before the build, a child holding state (a hash table, OPTIONAL buckets) kept it through the inner build: NOT EXISTS after OPTIONAL rose from 106 to 139 MiB peak with no annotations involved. The planner now passes estimates of the outer block and the body. An outer side estimated small against a triples-only body is seeded a chunk of up to 1,024 keys at a time, so the cost follows the outer side; if the seeded keys outgrow the estimate, it switches to one unseeded build. Otherwise the unseeded build runs before the child is opened, as on main. Buffered outer rows and each chunk's keys are charged to the query budget and released as they drain. The outer estimate treats a term decomposition as the join it is rather than a scan of every term.
Export reads the ledger's live links up front. It cloned every link's term into the edge map, then cloned every edge key again to group the asserted check by (subject, predicate). Terms are now moved into the map and the grouping sorts borrowed keys; only edges found unasserted are cloned. 200k annotated edges, Turtle to a sink (debug build): peak RSS 395 to 287 MiB, against 56 MiB with --raw-reifies. The set is still held whole.
An annotated edge's expansion ends in an existence check on the base edge,
which the OPTIONAL lane's safety check rejected, so OPTIONAL { s p o {| … |} }
and a Cypher relationship binding went back to per-row rebuilds. An EXISTS of
plain triples reads only the row's bindings, as a triple does, and is now
admitted. It also counts as tolerating an unbound variable, so a row whose
correlation key is unbound still takes the per-row path.
The arena's annotation_planner bench went with the arena, leaving no bench on the lanes this branch adds. query_hot_annotation_links covers the base-edge EXISTS below and past one seeded build's 1,024 keys over a much larger predicate (constant and variable predicate), a full reified-pattern count, the batched wildcard join and the COUNT product, and checks each scenario's count once before timing. annotation_varlength_probe's doc now describes the link expansion instead of the retired sidecar drain.
…ey bundle refs Re-asserting a deleted edge brings its earlier claims back, since the link names the triple: say so, and how to remove them. The design doc now states how the lanes keep term policy and which reads a pre-link index refuses. Comments citing the removed EdgeKey bundle decoder name fold_slots_into_links instead.
|
There was an error handling pipeline event c9017ac8-e6ef-45a5-b91c-8af3b9772987. |
Export read every live link into an edge map before writing a line, which grows with the ledger. Up to 1M links (by the index's link counts) it still does, since a hash probe per edge is the fast path. Past that, or with no counts, each batch probes its own edges (a link lookup only for an edge whose term a dictionary holds) and checks each link row's triple, so memory follows the batch. The reifier bookkeeping keeps a hash per reifier instead of its IRI. 200k annotated edges (debug, Turtle to a sink): preloaded 287 MiB / 10.4 s, per-batch 133 MiB / 15.5 s, against 56 MiB / 2.0 s with --raw-reifies. Export also decoded a triple-term handle only through the persisted dictionary, so a term only novelty holds failed the export when written as a value, an unasserted link, or under --raw-reifies. It now falls back to dictionary novelty, as subjects and strings do.
|
There was an error handling pipeline event c90fe58c-e6ef-45a5-b91c-8af3b9772987. |
With no planner estimates the semijoin seeded chunk after chunk however large the outer side, rebuilding its partial-key lookup per chunk where it used to build the body once. It now seeds only when the first chunk holds the whole outer side and otherwise builds unseeded once, as before; with estimates, chunking is unchanged. optional_exists_reuses_partial_keys pins the lookup to once per seeded chunk at most, never per row.
|
There was an error handling pipeline event f445e8ae-5c0e-45cd-94f6-dde9a0af0219. |
A flake under a schema predicate (rdfs:range, rdfs:domain, ...) skipped the policy check, the triple-term check with it, so `?d rdfs:range ?t` returned a term naming a hidden triple, and OBJECT(?t) its value. The term is now checked first in both the batch and single-flake filters.
isLITERAL classed every non-resource as a literal, so it returned true for a triple term, and SPARQL DATATYPE returned f:tripleTerm from novelty and @id from an index. RDF 1.2 makes a triple term its own kind: isLITERAL is now false, SPARQL DATATYPE is a type error, and JSON-LD's datatype names f:tripleTerm on both paths, as it names @id for an IRI. The indexed literal count no longer counts TRIPLE_TERM rows.
The operator gathered every candidate of an input batch into one output batch, so decomposing a predicate's 2,500 terms with a batch size of 4 came back as one batch of 2,500, and none of the dictionary reads were charged. It now resumes mid-row across calls and emits at most a batch at a time, charges the dictionary reads it caches per input batch (released when the next batch starts) and the novelty term set, and checkpoints while it iterates. Novelty's terms are collected when a row first needs them, not on every open.
The bound read only the index's link counts, so an unindexed ledger counted as zero and preloaded however many links novelty held. Preloading now stops and falls back to per-batch probes as soon as the links it has read pass the bound, novelty's included.
The formatters and OBJECT()/TermComponents each turned a term's object into a binding, keeping its datatype or language tag, in their own copy.
Transactions, import and query values each split `{"@id": {"@id": s, p: o}}`
on their own and had drifted: values took any key but `@id` as the
predicate, and only import rejected an object node with properties.
`triple_term_parts` now checks the shape once (a subject, one predicate,
one object that is a reference, value or nested term); each surface keeps
its own lowering.
|
There was an error handling pipeline event ec12bfe1-3c8b-4e3f-b4c7-13d3415ad2d9. |
Without a binary store no lane reads raw rows, so the triple-term check declined fast paths for nothing and broke the gate's coverage test. It now answers no there; rdf:reifies still declines under a restricting policy, which the test now pins.
|
There was an error handling pipeline event 61f6c017-5d99-418f-91c1-fd8114e24cc7. |
|
Review-body items not tied to a line:
|
A ledger whose index was built before triple-term links surfaced only when a read of its annotations failed. Loading one now logs a warning, `fluree info` prints one (local, and tracked from the server's response), and `ledger-info` carries `"needs-link-reindex": true` in its `ledger` block. Pinned against the 4.2.3-written fixture.
|
There was an error handling pipeline event 3f4bb04a-957f-42ff-87de-f36d57356e6b. |
|
There was an error handling pipeline event 62bcf29f-99d0-4afb-980b-fdd45a53c8b1. |
|
@aaj3f on the release question: I've retargeted this PR to That moves the port you described (threading the object type through On the reindex itself: it only applies to ledgers that hold annotations and were indexed by an earlier release. Other ledgers, and the non-annotation reads of an annotated one, are unaffected. Such a ledger now says so up front: a warning on load, a line in Closing keywords don't fire on |
Summary
Edge annotations move to the RDF 1.2 model. Each reifier is stored as an ordinary
r rdf:reifies <<( s p o )>>link whose object is a triple term. That replaces the annotation arena and thef:reifies*bundles. Triple terms become first-class values under any predicate, nested ones included, on every read and write surface. The branch also reworks the annotation query paths, which more than halves total time on a 61M-triple annotation workload.Behavior changes
s p o {| … |},s p o ~ r) assertss p o.<< s p o >>andr rdf:reifies <<( s p o )>>reify it without asserting it. This holds on every write surface.rdf:reifieslinks are ordinary data. Wildcard reads (?r ?p ?o,<< … >> ?p ?o) return them, so counts over annotated ledgers change by the number of reifiers. Only the legacyf:reifies*predicates stay hidden.f:reifies*bundles, with no term dictionary. A reified-triple or annotation query against such an index fails with an error namingfluree reindex <ledger>, rather than answering as if there were no annotations. If such a ledger takes new annotation writes, its next index falls back to a full rebuild on its own. There is no in-place migration. Older releases can't read an index that carries a term dictionary.What's new
ex:doc ex:mentions <<( s p o )>>can be stored, matched, filtered and returned. This covers SPARQL, JSON-LD, Turtle, TriG, N-Triples, N-Quads and bulk import, nested terms included.TRIPLE(),SUBJECT(),PREDICATE(),OBJECT()andisTRIPLE(),=on triple terms, and triple terms inVALUES. The JSON-LD query surface has twins of all of these.INSERT DATA,DELETE DATA,DELETE WHEREand templates.{ … }default-graph block and unescapes local names inside graph blocks. Previouslyex:a\-bin aGRAPHblock was stored with its backslash.�/�escaped.Performance
These are medians over two runs on the same 32-vCPU x86 host, comparing
mainat7e8dba0d2with this branch. The first workload is a 61M-triple biomedical graph with 21M reifiers, queried through 56 queries covering point lookups, selective joins and full scans. Total time for all 56 went from 1,436 s to 731 s.main?d ?e << ?s ?p ?o >>FILTER<< ?s ex:treats ?o >> ?p ?x<< ?s ?p ?o >> ?d ?e(now 55M rows, since therdf:reifieslinks are included)<< ?s ?p ex:C >> ex:src ?xFor example, this query counts the provenance records of every statement made about one concept:
It went from 149 ms to 60 ms in the table above, which includes a cold first run. Warm, it takes about 6 ms. Before the leaflet skip below, each probe walked every leaflet across the predicate's subject range.
The changes that matter most:
Regressions
LDBC SNB SF0.1 interactive reads (median of five cases each):
mainThe rest are within noise.
IC7 is 2.1× slower than
main. Its likes carry annotations, and with more than 1,024 distinct outer keys the existence check builds over the wholeLIKESpredicate. The arena answered that per edge. A seeded probe for large outer sides would close the gap; it is not in this PR.A few multi-second scan-heavy queries on the first workload are 6–17% slower (for example 5.9 s → 7.1 s), which is consistent with the extra
rdf:reifiesrows now in wildcard reads.Conformance
[ … ],[]or collections (32 tests, TriG GRAPH blocks reject anonymous blank nodes [ … ] on /upsert and import #1930 / trig upsert with anonymous blank nodes causes parse error #1511).rdf:dirLangString) and annotation-of-annotation remain unsupported, as before.Testing
testsuite-sparql.Fixes #1848
Fixes #1878
Fixes #1879
Fixes #1880
Fixes #1882
Follow-up: #1930
Follow-up: #1511
#1848 asked to unify two constructions of the
f:reifies*bundle; neither exists any more. #1878, #1879 and #1880 concern annotation lane selection, and #1882 the arena's seal state; the arena and its lanes are removed.