Skip to content

Fix word alignment when tokenization drops input words - #394

Merged
Ingvarstep merged 1 commit into
urchade:mainfrom
issacchan26:fix/preserve-word-indices
Sep 23, 2026
Merged

Ingvarstep merged 1 commit into
urchade:mainfrom
issacchan26:fix/preserve-word-indices

Conversation

@issacchan26

Copy link
Copy Markdown
Contributor

Problem

When a pretokenized word produces no subtokens (for example, an empty string, newline, tab, or spaces), prepare_word_mask assigns consecutive positions to the remaining words. Span annotations still refer to the original input positions, so the model can silently receive supervision for the wrong word.

For example, with the tokens "Alice", a standalone newline, "joined", "Acme", and "yesterday", an Organization annotation at zero-based position 3 should supervise "Acme". Previously, the representation at position 3 belonged to "yesterday".

Change

Build the word mask from the tokenizer's original word IDs, subtracting the prompt length and converting to one-based mask indices. Tokens without subtokens retain a gap in the word positions. Keep the existing first/last/mean/max subtoken selection and token-level behavior.

The regression tests use a local WordPiece tokenizer and a tiny randomly initialized model. They require no network access or pretrained weights and cover:

  • Leading, interior, trailing, and consecutive invisible tokens.
  • Missing prompt words, different prompt lengths, and batch padding.
  • All four pooling modes and the token-level override.
  • Original span targets, zero embeddings at missing-word positions, and gradients reaching the annotated word.
  • Finite model loss and gradients.

Validation

  • New regression tests on unmodified main: 29 failed, 4 passed. With the fix: all 33 passed.
  • python -m pytest -q --tb=short --deselect tests/test_models.py::test_span_model: 591 passed, 4 skipped, 1 deselected. The deselected test requires downloading a pretrained GLiNER checkpoint; the other tests ran offline after caching public tokenizer/configuration files.
  • ruff check gliner tests/test_word_alignment.py: passed.
  • ruff format --check tests/test_word_alignment.py: passed.
  • Confirmed the generic example also keeps its original positions with the DeBERTa v3 fast tokenizer.

Test environment: Python 3.12, PyTorch 2.14.0, Transformers 4.57.6, Tokenizers 0.22.2. Modified files also parse with Python 3.10 grammar.

@Ingvarstep

Copy link
Copy Markdown
Collaborator

Thanks for fixing it, @issacchan26 !

@Ingvarstep
Ingvarstep merged commit 5803656 into urchade:main Sep 23, 2026
3 checks passed
Sign up for free to join this conversation on GitHub. Already have an account? Sign in to comment

Labels

None yet

Projects

None yet

Development

Successfully merging this pull request may close these issues.

2 participants