Skip to content

RDF 1.2 triple terms: rdf:reifies links, triple terms as values, faster annotation queries - #2017

Merged
bplatz merged 94 commits into
nextfrom
feat/annotation-triple-terms
Oct 6, 2026
Merged

bplatz merged 94 commits into
nextfrom
feat/annotation-triple-terms

Conversation

@bplatz

@bplatz bplatz commented Oct 5, 2026 •

Copy link
Copy Markdown
Contributor

Summary

Edge annotations move to the RDF 1.2 model. Each reifier is stored as an ordinary r rdf:reifies <<( s p o )>> link whose object is a triple term. That replaces the annotation arena and the f:reifies* bundles. Triple terms become first-class values under any predicate, nested ones included, on every read and write surface. The branch also reworks the annotation query paths, which more than halves total time on a 61M-triple annotation workload.

Behavior changes

  • A reified triple no longer asserts its triple. As RDF 1.2 defines it, only the annotation syntax (s p o {| … |}, s p o ~ r) asserts s p o. << s p o >> and r rdf:reifies <<( s p o )>> reify it without asserting it. This holds on every write surface.
  • rdf:reifies links are ordinary data. Wildcard reads (?r ?p ?o, << … >> ?p ?o) return them, so counts over annotated ledgers change by the number of reifiers. Only the legacy f:reifies* predicates stay hidden.
  • Deleting a triple leaves its reifiers in place, except in LPG edge-lifecycle mode.
  • Reindex required for ledgers that already hold annotations. Existing indexes load as they are. Ledgers without annotations need nothing. In an annotated ledger indexed before this branch, the annotations exist only as f:reifies* bundles, with no term dictionary. A reified-triple or annotation query against such an index fails with an error naming fluree reindex <ledger>, rather than answering as if there were no annotations. If such a ledger takes new annotation writes, its next index falls back to a full rebuild on its own. There is no in-place migration. Older releases can't read an index that carries a term dictionary.

What's new

  • Triple terms as values. ex:doc ex:mentions <<( s p o )>> can be stored, matched, filtered and returned. This covers SPARQL, JSON-LD, Turtle, TriG, N-Triples, N-Quads and bulk import, nested terms included.
  • SPARQL 1.2 expressions. TRIPLE(), SUBJECT(), PREDICATE(), OBJECT() and isTRIPLE(), = on triple terms, and triple terms in VALUES. The JSON-LD query surface has twins of all of these.
  • SPARQL UPDATE takes reified triples, named-graph annotations and triple terms in INSERT DATA, DELETE DATA, DELETE WHERE and templates.
  • CONSTRUCT templates write reified triples, annotations, and triple terms under any predicate, constant or built from each solution, nested ones included. This works in both SPARQL and JSON-LD templates.
  • TriG ingest accepts the unlabeled { … } default-graph block and unescapes local names inside graph blocks. Previously ex:a\-b in a GRAPH block was stored with its backslash.
  • The N-Triples and N-Quads writers emit canonical form: lowercase language tags, and � / � escaped.

Performance

These are medians over two runs on the same 32-vCPU x86 host, comparing main at 7e8dba0d2 with this branch. The first workload is a 61M-triple biomedical graph with 21M reifiers, queried through 56 queries covering point lookups, selective joins and full scans. Total time for all 56 went from 1,436 s to 731 s.

Query shape main this branch
Count of an annotated-edge chain joined with an independent full predicate scan 214 s 0.27 s
Any edge whose object is a reifier: ?d ?e << ?s ?p ?o >> 156 s 10.5 s
Join of two annotated-edge groups with constant anchors, compared by a FILTER 34.7 s 5.7 s
Count of every annotation on edges with one predicate: << ?s ex:treats ?o >> ?p ?x 7.2 s 1.2 s
Every annotated edge with every annotation: << ?s ?p ?o >> ?d ?e (now 55M rows, since the rdf:reifies links are included) 414 s 128 s
Annotations of every edge into one node: << ?s ?p ex:C >> ex:src ?x 149 ms 60 ms
Annotations of two constant edges, cross-joined 30 ms 12 ms

For example, this query counts the provenance records of every statement made about one concept:

SELECT (COUNT(*) AS ?n) WHERE {
  << ?s ?p ex:someConcept >> ex:derivedFrom ?source .
}

It went from 149 ms to 60 ms in the table above, which includes a cold first run. Warm, it takes about 6 ms. Before the leaflet skip below, each probe walked every leaflet across the predicate's subject range.

The changes that matter most:

  • Annotated edges check their base triple with a semijoin, not a join. The semijoin is seeded from the outer side's own keys when that side is small.
  • Link probes bound by a triple-term handle use the bulk object-probe lane, instead of decoding the term and re-deriving its handle per row.
  • Batched subject probes reject PSOT leaflets on their directory keys before loading them.
  • Wildcard-predicate joins use a shared streaming lookup.
  • An independent right-hand scan is read once and replayed. COUNT over a Cartesian product of independent sides is computed as a product.

Regressions

LDBC SNB SF0.1 interactive reads (median of five cases each):

Query main this branch
IC5 2,040 ms 1,095 ms
IC7 67 ms 141 ms
IC1 41 ms 46 ms

The rest are within noise.

IC7 is 2.1× slower than main. Its likes carry annotations, and with more than 1,024 distinct outer keys the existence check builds over the whole LIKES predicate. The arena answered that per edge. A seeded probe for large outer sides would close the gap; it is not in this PR.

A few multi-second scan-heavy queries on the first workload are 6–17% slower (for example 5.9 s → 7.1 s), which is consistent with the extra rdf:reifies rows now in wildcard reads.

Conformance

Testing

  • New paths: each has a test, and each test was checked to fail with its fix reverted.
  • Workspace gates: fmt and clippy pass on the workspace and on testsuite-sparql.
  • Suites: the core, graph, index, transact, sparql, query, CLI and full API suites pass, as do all W3C SPARQL and RDF suites.

Fixes #1848
Fixes #1878
Fixes #1879
Fixes #1880
Fixes #1882
Follow-up: #1930
Follow-up: #1511

#1848 asked to unify two constructions of the f:reifies* bundle; neither exists any more. #1878, #1879 and #1880 concern annotation lane selection, and #1882 the arena's seal state; the arena and its lanes are removed.

bplatz added 30 commits October 2, 2026 09:04
OPST holds every object type, segmented by o_type. IndexType::for_query gates it on o_is_ref only as a conservative default, and the binary scan path already takes OPST for any constant object.
A triple term `<<( s p o )>>` becomes a first-class object value.
`FlakeValue::TripleTerm` carries the materialized form; the index stores
a dictionary handle under `OType::TRIPLE_TERM` (tag-10, decode kind
`TripleTermDict`, V1 kind byte 0x15). Handles are partitioned by the
inner predicate (`triple_term::term_handle`) so a predicate restriction
is one `o_key` interval, and `TermKey` is the encoded base edge, ordered
subject-first for the reverse tree.

Adds the `rdf:reifies` and `f:tripleTerm` sids, plus a read-side
`is_scan_hidden_predicate` so wildcard scans keep hiding the link form
alongside the `f:reifies*` bundle while it is index-internal. Every
closed match over `FlakeValue` gains an arm: formatters render the term
by its display form for now, and the commit codec rejects the value
until the write path moves over.
The link `_:r rdf:reifies <<( s p o )>>` is now one main-index flake on
imported ledgers, alongside the `f:reifies*` bundle it is meant to
replace. Its object is a handle into a ledger-global term dictionary
that mirrors the subject dictionary: one forward pack stream per inner
predicate keyed by sequence, one reverse tree keyed by the encoded base
edge, and references in a flagged trailing root section
(`FLAG_EXT_HAS_TERM_DICT`) that the core metadata parser skips.

Import interns terms after the global remap, where ids are final. The
sink records each reified edge as a pseudo-record in a per-chunk term
table and spools the link with the entry's ordinal; both build phases
resolve the ordinal through `TermRemapCtx`, which remaps the entry with
the chunk's own tables and interns it in a build-wide `TermDictBuilder`.
The dictionary is uploaded next to the index artifacts and referenced
from the root, and the store decodes a handle back to the materialized
term.

Full rebuilds and incremental builds do not intern yet (a rebuild drops
the term section), and the query lowering still takes the bundle path.
…builds

The resolver now assembles each commit's `f:reifies*` bundle into the
RDF 1.2 link record and a per-chunk term entry (`link_synth`), so an
index built from commits carries the same `rdf:reifies` flakes bulk
import writes. This is what makes the CLI's post-import seal, and every
later reindex, keep the links and the dictionary.

Full rebuilds intern in Phase C once ids are global and publish the
dictionary from the root. Incremental builds look each term up in the
base dictionary, allocate fresh handles above its per-predicate
watermarks, append per-predicate forward packs and update the reverse
tree copy-on-write, with the replaced tree blobs recorded as garbage.
Pack appends take the pack kind explicitly, since the subject appender
hardcoded it.

The stage cascade drops the index-side link before partitioning a
reifier's flakes into bundle and body, so an indexed reifier whose body
is deleted still looks orphaned and cascades; the link itself is retired
by the next index pass. Tests cover import, reindex and the incremental
window, including a second reifier on an existing edge.
Behind `FLUREE_ANNOTATION_TERMS=1`, the SPARQL lowering routes the
reified-triple forms (`<< s p o >>` and `?r rdf:reifies <<( s p o )>>`)
through the link flake instead of the `f:reifies*` bundle chain: one
`annotation rdf:reifies ?__term_N` triple, then `BIND(SUBJECT(?__term_N)
AS ?s)` for a component variable no earlier pattern binds, or a `FILTER`
equality when one does, since `BIND` overwrites. A constant predicate
becomes `FILTER(PREDICATE(?__term_N) = <p>)`, which the filter pushdown
turns into `ObjectBounds::term_predicate` and the scan into one `POST`
handle interval: the seek keys narrow the leaflets, and each row's
`o_key` is checked against the interval in encoded form, so a count
never decodes a term. A fully constant edge composes to a constant term
the scan resolves through the reverse tree.

`SUBJECT`, `PREDICATE`, `OBJECT` and `isTRIPLE` now evaluate over
`FlakeValue::TripleTerm`; `TRIPLE()` stays deferred. Link scans emit
late-materialized handles. The annotation syntax `s p o {| |}` and the
JSON-LD surface still take the bundle chain.

The link tests move to their own binary because the flag is
process-wide; the new test pins every access-path shape from the design
doc, including a component variable bound by an earlier pattern.
… the link lowering

The flagged link lowering was slower than the bundle chain: a composed
constant term never reached the dictionary (the simple encoder had no
TripleTerm arm, so a fully bound pattern fell to a decode-and-compare
scan of every link), every SUBJECT/PREDICATE/OBJECT accessor materialized
the whole term per row, and positions nothing reads were still bound.

- compose_term_handle is shared by both encoders; an uninterned term is
  NotFound, so the scan skips straight to novelty.
- An accessor on a late-materialized handle reads the dictionary key and
  binds the component encoded (EncodedSid, EncodedPid, EncodedLit with
  its datatype and tag); OBJECT on a materialized term keeps the
  datatype and tag too.
- elide_unread_term_binds drops component BINDs nothing reads, the twin
  of elide_redundant_chain for the bundle chain.
- Variable positions are always BIND: the agreement check is the join in
  every scope, and bind_unifies normalizes representations (a scan's
  EncodedSid against a VALUES row's Sid) in BindOperator and apply_inline.
  The lowering's cross-scope binding set is gone.
- Constant positions are sameTerm filters the pushdown turns into
  ObjectBounds::{term_subject, term_object} beside term_predicate; the
  scan checks them on the handle's dictionary key (term_key_filter). A
  literal object carries its datatype or tag as a resolved binding, so
  "5"^^xsd:int and "chat"@fr match by term identity, as the composed
  constant does. Two constraints on one component that disagree keep the
  second filter instead of one silently winning.
- Import skips arena-scoped (decimal, big-integer, vector) objects as
  rebuild does: per-(graph, predicate) handles are no graph-independent
  term identity.
…undles per commit

A re-point that changes one slot writes only that slot's retract and
assert (sync and upsert cancel the unchanged slots), and a full
re-point's six ops interleave under the per-commit sort, so an assembler
that needed all three slots in one (graph, reifier, t, op) group emitted
no link events and the old link survived the build.

The resolver now only collects f:reifies* ops (RebuildChunk.attachments)
and the build replays them per reifier once ids are global: ops in t
order, retracts before asserts within one t, a link retract and assert
wherever the attachment changes from or to a complete edge. A rebuild
replays the whole history from nothing and writes the links as one more
sorted commit file; an incremental build seeds each touched reifier with
the attachment the base index holds (one SPOT point lookup) and appends
the links to the window.
…s look

The count planner and the other fused fast paths classify the query's
pattern list as lowered, so a `COUNT(*)` over `<< ?s ?p ?o >>` still
carried three component BINDs at that point and fell through to the
generic join, ten times slower than the same two-triple join written by
hand. Run the unread-position elision on the query before any fast path
sees it; the WHERE planner's own pass stays for the seeded and nested
cases.
OBJECT() of a materialized term, and DATATYPE/LANG over any accessor, went
through the numeric comparable, which carries no datatype: an indexed
"5"^^xsd:int read as xsd:integer as soon as an unrelated unindexed commit
turned late materialization off, and DATATYPE(OBJECT(?t)) lost the subtype
even on a quiescent ledger.

term_component_binding now yields the component as a binding on both the
encoded path (dictionary key) and the materialized path (the term itself),
with datatype and tag, and DATATYPE and LANG read an accessor argument
through it exactly as they read a bound variable. Pinned on both paths.
…nk replay

The incremental replay read a reifier's base predicate slot with a
swallowed error, so a failed dictionary fetch looked like a missing
predicate, the replay saw an incomplete attachment before and after an
object-only re-point, and the old link survived. The read error and a
predicate absent from the dictionary now fail the build.

Rebuild frees each chunk's record and term buffers as soon as its spool
file is written instead of after the loop, and the replay hands link
records to the spool one reifier at a time, in (graph, SPOT) order, rather
than accumulating them beside the attachment table.
A reified-triple pattern lowers to `?r rdf:reifies ?t`, and a row such as
`ex:r rdf:reifies ex:x` satisfied it: the accessor BINDs failed and kept
the row, and a count elided them altogether. Neither a filter nor a
datatype constraint on the link triple is an option, since both take the
shape off the count plan and the planner paths that gate on dtc.is_none(),
and the count planner counts from leaf-directory metadata that cannot see
object types.

So the predicate carries the rule. The write path refuses rdf:reifies with
a non-term object on every surface (JSON-LD, SPARQL UPDATE, Turtle
transact, the template path, bulk import), as it refuses user-authored
f:reifies*; the reified-triple forms never reach those checks. The scan
treats an rdf:reifies pattern with a variable object as term-only on the
encoded object type, on novelty flakes too, and narrows a POST seek to the
term segment. Rows written as a plain predicate before this rule are
dropped by the scan and still counted by the count plan.
An incremental build replays each touched reifier from the attachment the
base index holds, and a re-point retracts the old term's link. Both lookups
ran once per reifier: a fresh SPOT cursor for the base attachment, and a
reverse-tree lookup for each term, which rereads a whole ~2.3 MB leaf when
no leaf cache is attached.

The base attachments now come from one batched SPOT pass per graph
(batched_lookup_subject_properties), with each predicate IRI resolved once.
Term handles are resolved together: a dry replay names the terms the real
replay will ask for, and TermDictReader::find_handles reads each leaf once.

On the 10% annotation benchmark slice, an incremental index over 100k partial re-points
drops from 72.7 s to 30.2 s (the bundle path on main: 27.0 s), at the same
peak memory.
`fluree_db_core::io_stats` counts artifact reads by kind and source, and
the CLI prints the table on exit when FLUREE_IO_STATS is set. Remote
fetches are counted once, at the storage bridge (`store` / `store-range`),
with dictionary blobs told apart by format (string, subject and term
packs, reverse-tree leaves and branches); readers count what a local path
or the disk artifact cache served.

FLUREE_FORCE_REMOTE_READS makes a local store answer as a remote one: no
resolved local paths and no resident bytes, so every artifact goes through
`get` / `get_range` and the disk cache, and the counts are what S3 would be
asked for.

Off, each hook is one relaxed atomic load and builds no label.
OType::TRIPLE_TERM joined the Fluree-reserved entries, so the built-in table holds 45 entries, not 44.
…al builds

Phase 3b seeds every class entry of the base stats and resolves each class
Sid to its sid64 through the subject reverse tree, one lookup per class.
Without a leaf cache each lookup rereads a whole ~2.3 MB leaf, so a ledger
with many classes paid it on every incremental pass: 203,269 classes took
21 s on the 10% annotation benchmark slice, whatever the size of the window.

`BinaryIndexStore::find_subject_ids_by_parts` resolves them together through
`reverse_lookup_many`, reading each leaf once. A failed read now aborts to a
full rebuild, as the base PSOT class lookup beside it does, instead of
dropping the class from the published stats.
Loading an indexed ledger read the root twice: `LedgerState::load` for the
snapshot metadata, then the binary store load for the full decode. Against
a remote store that is a second whole-object GET of the largest single
artifact on every cold load (83 MB twice on the annotation benchmark slice).

`LedgerState::load_with_root` / `load_with_store_and_root` hand back the
root they read, and `load_and_attach_binary_store` takes it when its CID is
the one the record names. Every load-then-attach path passes it through.
The root's stats carry per-graph class tables and a ledger-wide one, and
both build paths set the ledger-wide table to `union_per_graph_classes` of
the per-graph ones. On a ledger with many classes the two copies are most
of the root: 41.5 MB each for 203,269 classes on the annotation benchmark slice.

The encoder now writes the ledger-wide table only when no graph carries
classes, and both decoders derive it from the graphs when it is absent.
Roots that still carry it decode as before. Import roots, which never set
it, now present the same class table a rebuild would, instead of none until
the next reindex. A reader without this change sees no ledger-wide class
table on a new root and plans without class statistics.
Each class-property usage spelled out its property Sid, and each ref-class
count its class Sid, with fixed-width counts. On a ledger with many classes
that is nearly the whole root: 203,269 classes, 758,489 usages and 347,070
ref counts took 41.5 MB on the annotation benchmark slice, though only 607 predicates
appear among the usages.

The class tables now travel in an appended stats section (tag 2): every Sid
written once in a sorted, front-coded table, entries referring to it by
index, counts as varints. Re-encoding the slice's 83 MB root gives 8.4 MB
with identical decoded stats, and decoded entries share their Sid names.

Readers skip appended sections they do not know, so the legacy per-graph
class slots are left empty and a reader without this change sees no class
tables rather than misparsing. Roots that carry classes in the legacy slots
still decode. The binary-index stats decoder now delegates to the core one,
so the format has a single parser.
The state load decoded the root for the snapshot metadata
(`LedgerSnapshot::from_root_bytes`) and the binary store load decoded the
same bytes again as an `IndexRoot`; on the annotation benchmark slice each decode took
about 110 ms.

`LedgerState::load_decoding_root` / `load_with_store_decoding_root` take the
root decoder from the caller and hand back what it kept. The API decodes the
root once as an `IndexRoot`, builds the snapshot from it through
`IndexRoot::take_snapshot_metadata` (stats and schema move into the snapshot
rather than being copied), and passes the rest of the root to the binary
store load. Index publish builds its snapshot through the same method.
`load` and `load_with_store` keep the metadata-only decode.
Each incremental build appended a pack to every inner predicate's term
stream that gained terms, and nothing merged them, so a predicate's routing
table grew by one entry per build forever. The streams now go through the
same tail-first compaction as the subject and string streams, within the
cycle's shared budget (hoisted so every stream draws on one allowance).
Packs a merge consumes, base packs and this cycle's own uploads alike, are
recorded as garbage with the term dictionary's replaced reverse-tree leaves.
…aries

The server's background warmer touched string and subject packs only, so
the first query to decode reified edges fetched every term pack it needed
one at a time. Term packs are now warmed last, within the same budget.
A BIND into a variable that an earlier step already bound is an equality
check against that column. When the bind's expression waits on a later
pattern, the planner holds it back, and each intermediate join step trims
its output to the variables still live: later triples, pending filters, and
the variables pending binds' expressions read. The bind's own target was
not among them, so a target nothing else read was trimmed, the bind found
no value and bound afresh, and the check disappeared: a count became the
cross product.

JSON-LD reaches it with a rebind (`["bind", "?a", "?v"]` after `?a` is
bound and one more join): 6 rows instead of 1. The reified-edge lowering
reaches it from SPARQL, since every quoted position is a BIND and a shared
component variable is joined that way: `<< ex:a ex:p ?o1 >> ... << ?o1 ex:q
?o2 >> ...` under COUNT(*) counted 167,551,488 instead of 189 on the
annotation benchmark slice. A pending bind now keeps its target live.
…ing templates

The write firewall checked `rdf:reifies` where a surface spells the
predicate out. A predicate variable is resolved only when a template is
materialized per solution row, so `INSERT { ex:r ?p ex:x } WHERE { VALUES
?p { rdf:reifies } }`, or the JSON-LD update with a `?p` key, wrote a
non-term link row. The count plan answers from leaf metadata, which would
count it, while the scan drops it.

Template materialization now refuses an asserted `rdf:reifies` flake whose
resolved object is not a triple term, which covers SPARQL UPDATE, JSON-LD
update and every other template path. Retracting a legacy row still works.
`term_component_binding` produced a binding, with the object's datatype or
tag, only when the accessor's argument was a bare variable. Any other
argument (`OBJECT(COALESCE(?t))`) fell through to the numeric comparable,
which carries no datatype, so `DATATYPE` reported `xsd:integer` for an
`xsd:int` object and a `BIND` over the accessor lost the subtype too.

The general path now evaluates the argument to a term and builds the same
binding; the encoded-handle shortcut for a bare variable is unchanged, and
the comparable accessor converts that binding like any other.
…weep

Term streams were compacted only in a cycle that added terms to them, after
the string and subject streams had drawn on the shared budget, so a term
stream that missed its turn and then went quiet stayed fragmented for good.

The term dictionary now reserves a third of each cycle's compaction budget
before the other dictionaries run, and takes over what they leave. Its
streams that grew are compacted first, then the rest in a rotation keyed by
the base index t, as the subject maintenance sweep does; a cycle with no
new terms still runs the sweep and publishes the merged routing table.
…later one

`DictTreeReader::reverse_range_scan` found its first leaf with `find_leaf`,
which answers only for a key inside some leaf's [first, last] range. A start
key in the gap between two leaves (past one leaf's last key, before the
next one's first) found no leaf, and the scan returned nothing although the
later leaves held matching keys. A subject prefix lookup
(`find_subjects_by_prefix`) whose prefix sorts into such a gap missed every
match.

The scan now starts at the first leaf whose last key is not below the start
key.
…tern

The link lowering decomposed `<< s p o >>` into `?r rdf:reifies ?t` plus a
BIND per variable position. Component equality was then visible only after
the link scan: `<< X :treats ?o1 >> ... << ?o1 :causes ?o2 >> ...` read every
`:causes` link once per left row and joined on `?o1` in the BIND check, 12 s
against 90 ms for the bundle chain on the annotation benchmark slice.

Quoted patterns now lower to the link plus `TermComponents(?t, s, p, o)`,
an internal pattern the planner places like any other source. Its operator
finds candidate terms per input row:

- the term variable bound: one dictionary key decode supplies every
  component;
- the subject bound, as a constant or by an earlier pattern: the reverse
  tree's subject-first `(s_id[, p_id])` prefix range, which yields keys and
  handles together; distinct prefixes are read once per batch;
- neither: the terms of the bound predicate, or of the dictionary.

A variable component another pattern bound first is joined on with the
same equality a BIND check applied. The planner prices a bound term as a
decode and an anchored subject as a small source, so the chained edge is
found through its subject and its link becomes a bound-object lookup; the
link keeps the graph, time, policy and novelty checks. Constant components
stay filters on the link as well, so a link-first plan keeps its handle
interval. Unread positions are elided as before, and a pattern left with
nothing to bind or anchor goes, which keeps the count plan.

On the slice the populated chained query takes 60-75 ms (bundle chain 92 ms),
with every other measured shape unchanged or faster.
… covered

Links existed only in the index, so a link query missed every annotation
added, retracted or re-pointed since the last build. Novelty now derives
them from each commit's f:reifies* slot ops, replaying a reifier from the
index's attachment plus its earlier novelty ops, the rule the index build
applies. The index's attachments come from a base the ledger installs: an
empty one when the index never held an annotation, the binary store's
otherwise (at load, at each publish, for historical views). Commits applied
before the base is known wait and are derived when it arrives; a trim past
the base drops it.

Terms only novelty names take the raw-flake lane, and TermComponents offers
them beside the dictionary's, also on a ledger with no index. Graph
management and the lifecycle probe skip derived links: no transaction
writes one.
A link whose term the index has not interned took the raw-flake lane, so
any annotation committed since the last build took a link count off the
count plan (P2 on the dev slice: 25 ms indexed, 2.8 s with 10k such
links). Dictionary novelty now keeps a term table, populated from the
links novelty derives, and a term there gets a provisional handle in its
inner predicate's interval, above every indexed sequence: 2^31 plus its
index. The term-dictionary builder stops below that.

Composition tries the persisted dictionary first. A provisional handle's
key is computed from its novelty term, the novelty-aware graph view and
the batched join decode it, and publish retires the entries the index now
interns. apply_commit returns the links it derived so each write path
registers them. P2 with the same novelty: 69 ms, on the count plan.
A scan whose bound object was a term walked every novelty flake of its
predicate, so each chained link lookup re-translated all of novelty's
links, composing each term (P19b over 10k unindexed annotations: 12 s).
Terms order by base subject and predicate first, so the terms sharing the
bound term's form one run, which the walk now brackets.
bplatz added 8 commits October 6, 2026 07:52
A triple-term object is checked as the triple it names, but the scan and
probe lanes skip per-flake filtering whenever the scanned predicate is
uncovered by the view policy. On an indexed ledger a policy hiding
ex:salary therefore still returned the hidden triple through
`?r rdf:reifies ?t`, `<< ?s ex:salary ?o >> ...` and any other predicate
holding a term of it.

Under a view policy that restricts anything, the shared predicate gate now
declines `rdf:reifies` and any predicate whose observed datatypes include
the untagged kind triple terms carry, are unknown, or miss novelty the
stats cannot see. The policy test now runs from novelty, from an index and
from novelty over one, and covers a term held under another predicate.
On an annotated ledger indexed by an earlier release, link queries refused
with the reindex error, but export returned Ok with every reifier body
orphaned and no markers, and a crawl dropped @annotation: both read links
the index does not have.

The snapshot now carries `needs_link_reindex` (annotations, no term
dictionary), which also gives embedding hosts a machine-readable signal to
schedule the rebuild. Queries, export and hydration's annotation reads all
refuse through `require_link_index`, whose message now names the API route
as well as the CLI one. A ledger written and indexed by 4.2.3 is checked in
as a fixture to pin it.
…er side

Past 1,024 distinct outer keys the semijoin fell back to an unseeded build
over the whole body: an annotated edge's EXISTS on a large predicate cost
the predicate, not the query (IC7, and 1,100 hub likes over 1M ex:likes at
2-3x main's time and ~11x its memory). And because the outer side was
pulled before the build, a child holding state (a hash table, OPTIONAL
buckets) kept it through the inner build: NOT EXISTS after OPTIONAL rose
from 106 to 139 MiB peak with no annotations involved.

The planner now passes estimates of the outer block and the body. An outer
side estimated small against a triples-only body is seeded a chunk of up to
1,024 keys at a time, so the cost follows the outer side; if the seeded keys
outgrow the estimate, it switches to one unseeded build. Otherwise the
unseeded build runs before the child is opened, as on main. Buffered outer
rows and each chunk's keys are charged to the query budget and released as
they drain. The outer estimate treats a term decomposition as the join it is
rather than a scan of every term.
Export reads the ledger's live links up front. It cloned every link's term
into the edge map, then cloned every edge key again to group the asserted
check by (subject, predicate). Terms are now moved into the map and the
grouping sorts borrowed keys; only edges found unasserted are cloned.

200k annotated edges, Turtle to a sink (debug build): peak RSS 395 to 287
MiB, against 56 MiB with --raw-reifies. The set is still held whole.
An annotated edge's expansion ends in an existence check on the base edge,
which the OPTIONAL lane's safety check rejected, so OPTIONAL { s p o {| … |} }
and a Cypher relationship binding went back to per-row rebuilds. An EXISTS of
plain triples reads only the row's bindings, as a triple does, and is now
admitted. It also counts as tolerating an unbound variable, so a row whose
correlation key is unbound still takes the per-row path.
`{"@id": {"@id": s, p: o}}` in a values cell is a triple term, the twin of
SPARQL's `VALUES ?t { <<( s p o )>> }`. The object may be an IRI, a literal
or a nested term, built as the SPARQL cell and `triple()` build it.
The arena's annotation_planner bench went with the arena, leaving no bench
on the lanes this branch adds. query_hot_annotation_links covers the
base-edge EXISTS below and past one seeded build's 1,024 keys over a much
larger predicate (constant and variable predicate), a full reified-pattern
count, the batched wildcard join and the COUNT product, and checks each
scenario's count once before timing. annotation_varlength_probe's doc now
describes the link expansion instead of the retired sidecar drain.
…ey bundle refs

Re-asserting a deleted edge brings its earlier claims back, since the link
names the triple: say so, and how to remove them. The design doc now states
how the lanes keep term policy and which reads a pre-link index refuses.
Comments citing the removed EdgeKey bundle decoder name
fold_slots_into_links instead.
@azure-pipelines

Copy link
Copy Markdown
There was an error handling pipeline event c9017ac8-e6ef-45a5-b91c-8af3b9772987.

Export read every live link into an edge map before writing a line, which
grows with the ledger. Up to 1M links (by the index's link counts) it still
does, since a hash probe per edge is the fast path. Past that, or with no
counts, each batch probes its own edges (a link lookup only for an edge
whose term a dictionary holds) and checks each link row's triple, so
memory follows the batch. The reifier bookkeeping keeps a hash per reifier
instead of its IRI.

200k annotated edges (debug, Turtle to a sink): preloaded 287 MiB / 10.4 s,
per-batch 133 MiB / 15.5 s, against 56 MiB / 2.0 s with --raw-reifies.

Export also decoded a triple-term handle only through the persisted
dictionary, so a term only novelty holds failed the export when written as
a value, an unasserted link, or under --raw-reifies. It now falls back to
dictionary novelty, as subjects and strings do.
@azure-pipelines

Copy link
Copy Markdown
There was an error handling pipeline event c90fe58c-e6ef-45a5-b91c-8af3b9772987.

With no planner estimates the semijoin seeded chunk after chunk however
large the outer side, rebuilding its partial-key lookup per chunk where it
used to build the body once. It now seeds only when the first chunk holds
the whole outer side and otherwise builds unseeded once, as before; with
estimates, chunking is unchanged. optional_exists_reuses_partial_keys pins
the lookup to once per seeded chunk at most, never per row.
@azure-pipelines

Copy link
Copy Markdown
There was an error handling pipeline event f445e8ae-5c0e-45cd-94f6-dde9a0af0219.

bplatz added 6 commits October 6, 2026 09:03
A flake under a schema predicate (rdfs:range, rdfs:domain, ...) skipped the
policy check, the triple-term check with it, so `?d rdfs:range ?t` returned a
term naming a hidden triple, and OBJECT(?t) its value. The term is now
checked first in both the batch and single-flake filters.
isLITERAL classed every non-resource as a literal, so it returned true for
a triple term, and SPARQL DATATYPE returned f:tripleTerm from novelty and
@id from an index. RDF 1.2 makes a triple term its own kind: isLITERAL is
now false, SPARQL DATATYPE is a type error, and JSON-LD's datatype names
f:tripleTerm on both paths, as it names @id for an IRI. The indexed literal
count no longer counts TRIPLE_TERM rows.
The operator gathered every candidate of an input batch into one output
batch, so decomposing a predicate's 2,500 terms with a batch size of 4 came
back as one batch of 2,500, and none of the dictionary reads were charged.
It now resumes mid-row across calls and emits at most a batch at a time,
charges the dictionary reads it caches per input batch (released when the
next batch starts) and the novelty term set, and checkpoints while it
iterates. Novelty's terms are collected when a row first needs them, not on
every open.
The bound read only the index's link counts, so an unindexed ledger counted
as zero and preloaded however many links novelty held. Preloading now stops
and falls back to per-batch probes as soon as the links it has read pass the
bound, novelty's included.
The formatters and OBJECT()/TermComponents each turned a term's object into
a binding, keeping its datatype or language tag, in their own copy.
Transactions, import and query values each split `{"@id": {"@id": s, p: o}}`
on their own and had drifted: values took any key but `@id` as the
predicate, and only import rejected an object node with properties.
`triple_term_parts` now checks the shape once (a subject, one predicate,
one object that is a reference, value or nested term); each surface keeps
its own lowering.
@azure-pipelines

Copy link
Copy Markdown
There was an error handling pipeline event ec12bfe1-3c8b-4e3f-b4c7-13d3415ad2d9.

Without a binary store no lane reads raw rows, so the triple-term check
declined fast paths for nothing and broke the gate's coverage test. It now
answers no there; rdf:reifies still declines under a restricting policy,
which the test now pins.
@azure-pipelines

Copy link
Copy Markdown
There was an error handling pipeline event 61f6c017-5d99-418f-91c1-fd8114e24cc7.

@bplatz

bplatz commented Oct 6, 2026

Copy link
Copy Markdown
Contributor Author

Review-body items not tied to a line:

  • Stale bundle comments: addressed in 2b23f9e (where_plan.rs), da463d7 (edge_annotations.rs) and 18040ba (annotation_varlength_probe.rs).
  • Index version question: addressed in c7a5977. A loaded ledger exposes LedgerSnapshot::needs_link_reindex, and the error names Fluree::reindex alongside the CLI command.
  • The two extra behavior changes (--raw-reifies, escaped TriG local names) go into the description with the IC7 re-measurement.

A ledger whose index was built before triple-term links surfaced only when
a read of its annotations failed. Loading one now logs a warning,
`fluree info` prints one (local, and tracked from the server's response),
and `ledger-info` carries `"needs-link-reindex": true` in its `ledger`
block. Pinned against the 4.2.3-written fixture.
@azure-pipelines

Copy link
Copy Markdown
There was an error handling pipeline event 3f4bb04a-957f-42ff-87de-f36d57356e6b.

@bplatz
bplatz changed the base branch from main to next October 6, 2026 14:02
@azure-pipelines

Copy link
Copy Markdown
There was an error handling pipeline event 62bcf29f-99d0-4afb-980b-fdd45a53c8b1.

@bplatz

bplatz commented Oct 6, 2026 •

Copy link
Copy Markdown
Contributor Author

@aaj3f on the release question: I've retargeted this PR to next, so it lands with the upcoming major instead of on main. Nothing here holds up a release anymore. 4.2.4 with #2013 can go out from main whenever it's ready, and #2010 can land on main without waiting on this.

That moves the port you described (threading the object type through for_each_object_probe_match so term-handle probes keep working) to the merge of main into next. I'm happy to take that merge.

On the reindex itself: it only applies to ledgers that hold annotations and were indexed by an earlier release. Other ledgers, and the non-annotation reads of an annotated one, are unaffected. Such a ledger now says so up front: a warning on load, a line in fluree info, and needs-link-reindex in ledger-info (583c4db).

Closing keywords don't fire on next, so Fixes #1859 and the others stay in this description for the record and get carried into the next → main PR.

@bplatz
bplatz merged commit a25d446 into next Oct 6, 2026
16 checks passed
@bplatz
bplatz deleted the feat/annotation-triple-terms branch October 6, 2026 14:16
Sign up for free to join this conversation on GitHub. Already have an account? Sign in to comment

Labels

area:query Query execution, planning, fast paths, overlay, result formatting breaking-change Backwards-incompatible change; drives the Breaking Changes release-notes section required migration Requires updating file or index formats, ideal to package with other items requiring a migration.

Projects

None yet

2 participants