Skip to content
Open
Show file tree
Hide file tree
Changes from all commits
Commits
Show all changes
39 commits
Select commit Hold shift + click to select a range
7bea4c1
fix typos in docs
thorinaboenke Oct 8, 2026
ccf2801
improve wording
thorinaboenke Oct 8, 2026
e67bcec
correct hw to test docs examples
thorinaboenke Oct 8, 2026
9917096
typo
thorinaboenke Oct 8, 2026
9475bb6
parser schema and imports in example
thorinaboenke Oct 8, 2026
4d2b350
value range description
thorinaboenke Oct 8, 2026
9710763
remove duplication
thorinaboenke Oct 8, 2026
7893aaf
wording
thorinaboenke Oct 8, 2026
a2f5b11
correct example
thorinaboenke Oct 8, 2026
7165342
correct table
thorinaboenke Oct 8, 2026
dff2680
typo
thorinaboenke Oct 8, 2026
b87a70d
wording
thorinaboenke Oct 8, 2026
20d583d
add count vs row explanation
thorinaboenke Oct 8, 2026
04ed69b
replace arrows
thorinaboenke Oct 8, 2026
72c9ee9
correct file location
thorinaboenke Oct 8, 2026
7769a0a
remove stray line
thorinaboenke Oct 8, 2026
1a3cd5c
fix typos in federation example
thorinaboenke Oct 9, 2026
d04a6c0
correct claim about BigramFrequencyDetector
thorinaboenke Oct 9, 2026
e4e87e7
correct wording
thorinaboenke Oct 9, 2026
8eb48d6
make table match schemas.proto
thorinaboenke Oct 9, 2026
f79803f
correct logging example
thorinaboenke Oct 9, 2026
2d024ba
add info about event id
thorinaboenke Oct 9, 2026
9231f66
let tutorial use data independent from test fixtures
thorinaboenke Oct 9, 2026
2fccc7b
grammar
thorinaboenke Oct 9, 2026
e1a3261
remove pitfall from quickstart
thorinaboenke Oct 9, 2026
c36afcd
fix quickstart path problem
thorinaboenke Oct 9, 2026
fabe873
add navigation footer
thorinaboenke Oct 9, 2026
5d030bf
add requirements to installation
thorinaboenke Oct 9, 2026
3941820
change template file wording
thorinaboenke Oct 9, 2026
ab72e23
fix path confusion in tutorial
thorinaboenke Oct 9, 2026
c02689a
change pitfall wording
thorinaboenke Oct 9, 2026
cf7c419
schemas raise instead of returning exception class
thorinaboenke Oct 9, 2026
16874b3
corec class names
thorinaboenke Oct 9, 2026
839d04e
correct description of parse and run
thorinaboenke Oct 9, 2026
b1d6a8b
parser return value
thorinaboenke Oct 9, 2026
000b0cb
correct detector snippet!
thorinaboenke Oct 9, 2026
4e39a71
add axplanation for nemed eventid and autoconfig
thorinaboenke Oct 9, 2026
2dfef0b
remove double dashes
thorinaboenke Oct 9, 2026
55f95fb
corrections, add window behaviour
thorinaboenke Oct 9, 2026
File filter

Filter by extension

Filter by extension

Conversations
Failed to load comments.
Loading
Jump to
Jump to file
Failed to load files.
Loading
Diff view
Diff view
2 changes: 1 addition & 1 deletion README.md
Original file line number Diff line number Diff line change
Expand Up @@ -9,7 +9,7 @@ The library contains the next components:
* **Parsers**: parse the logs received from the reader.
* **Detectors**: return alerts if anomalies are detected.
* **Alert Aggregation**: aggregate the alerts produced by the detectors.
* **Schemas**: standard data classes use in DetectMate.
* **Schemas**: standard data classes used in DetectMate.
```
+--------+ +-----------+ +-------------------+
| Parser | --> | Detector | -> | Alert Aggregation |
Expand Down
4 changes: 1 addition & 3 deletions docs/advanced/overall_architecture.md
Original file line number Diff line number Diff line change
Expand Up @@ -25,7 +25,7 @@ Each arrow represents a stream of Schema objects. Components are designed to run

## Components architecture

All components inherit from a `CoreComponent` class. This class provides all the essential functionality required for DetectMate to operate (see UML diagram below). Every `Detector` must inherit from `CoreDetector`, every `AlertAggregator` must inherit from `CoreAlertAggregation` and every `Parser` must inherit from `CoreParser` to ensure compatibility with DetectMate.
All components inherit from a `CoreComponent` class. This class provides all the essential functionality required for DetectMate to operate (see UML diagram below). Every `Detector` must inherit from `CoreDetector`, every `AlertAggregator` must inherit from `CoreAlertAggregator` and every `Parser` must inherit from `CoreParser` to ensure compatibility with DetectMate.

Each component's arguments must be stored in its corresponding configuration class. These config classes follow the same design pattern as their components and must inherit from `CoreConfig`.

Expand All @@ -38,5 +38,3 @@ Each Core* base class exposes a small, stable API that implementations must impl
```python
--8<-- "docs/examples/others/components_methods.py:read"
```

Go back [Index](../index.md)
67 changes: 47 additions & 20 deletions docs/alert_aggregator.md
Original file line number Diff line number Diff line change
@@ -1,42 +1,69 @@
# Components: Alert Aggregation
# Alert Aggregation

Alert aggregation aggregates alerts from detectors.
Alert aggregators combine the alerts produced by detectors into aggregated records.

| | Schema | Description |
|------------|----------------------------|--------------------|
| **Input** | [DetectorSchema](schemas.md) | Alerts from detectors |
| **Output** | [AggregateSchema](schemas.md) | Aggregated alerts |

This document explains expected APIs, how to implement a parser, testing tips and common pitfalls.
This document explains the expected API and how to implement an alert aggregator.

## Overview

- Alert aggregation must inherit from `CoreAlertAggregation` and provide a `aggregate_alerts()` implementation.
- `CoreParser.run()` handles lifecycle and calls `aggregate_alerts()` for each input; implement pure alert aggregation logic inside `aggregate_alerts()` where possible.
- Use a typed `Config` class (subclass of `CoreAlertAggregationConfig`) to hold runtime parameters.
- Alert aggregators must inherit from `CoreAlertAggregator` and provide an `aggregate_alerts()` implementation.
- `CoreAlertAggregator.run()` handles the lifecycle and calls `aggregate_alerts()`. Aggregators use a
sliding window (a `WINDOW` [data buffer](auxiliar/input_buffer.md)) of `buffer_size` alerts, so
`aggregate_alerts()` always receives a **list**: the most recent `buffer_size` alerts. It is first
called once the window is full, and then once for every new alert, so consecutive windows overlap.
With the default `buffer_size` of 1, it is called for every alert with a list of one.
- Use a typed `Config` class (subclass of `CoreAlertAggregatorConfig`) to hold runtime parameters.

## CoreParser -- minimal API

Recommended signatures and behavior:
## CoreAlertAggregator: minimal API

```python
class CoreAlertAggregation:
def aggregate_alerts(
class CoreAlertAggregator(CoreComponent):
def __init__(
self,
name: str = "CoreAlertAggregator",
buffer_mode: BufferMode = BufferMode.WINDOW,
buffer_size: int | None = 1,
config: CoreAlertAggregatorConfig | dict[str, Any] | None = CoreAlertAggregatorConfig(),
) -> None:
"""buffer_mode and buffer_size set which alerts aggregate_alerts() receives per call."""

def aggregate_alerts(
self,
input_: list[DetectorSchema] | DetectorSchema,
output_: AggregateSchema,
) -> bool:
return True
"""Implement the aggregation here.
- Return True to emit output_ as the aggregated record.
- Return False to emit nothing: process() then returns None.
"""

@override
def train(
self, input_: DetectorSchema | list[DetectorSchema]
) -> None:
pass
def train(self, input_: DetectorSchema | list[DetectorSchema]) -> None:
"""Optional: learn from alerts. Can be a no-op."""
```

## Available alert aggregation
## What `run()` fills in for you

- [Basic Concat](alert_aggregators/basic_concatenation.md)
Before calling `aggregate_alerts()`, `CoreAlertAggregator.run()` collects these fields from the
alerts in the window:

Go back to [Index](index.md)
- `detectorIDs`, `detectorTypes`, `alertIDs`: one entry per alert
- `logIDs`, `extractedTimestamps`: the entries of all alerts, combined into one list

After `aggregate_alerts()` returns `True`, `run()` also sets `outputTimestamp`. So
`aggregate_alerts()` only has to set `description` and `alertsObtain`.

!!! warning "The alerts' details are not carried over"
`run()` copies only IDs and timestamps. The `description` and `alertsObtain` of the incoming
alerts, which hold the actual alert messages, are **not** copied into the aggregate, and
`AggregateSchema` has no `score` field at all. If your aggregator needs this information, read it
from `input_` in `aggregate_alerts()` and write it into `output_["description"]` or
`output_["alertsObtain"]` yourself.

## Available alert aggregators

- [Basic Concat](alert_aggregators/basic_concatenation.md)
4 changes: 1 addition & 3 deletions docs/auxiliar/input_buffer.md
Original file line number Diff line number Diff line change
@@ -1,6 +1,6 @@
# Data Buffer

The data buffer is an auxiliar methods that can be use in all the components. It takes the stream data and formated to the specifications given.
The data buffer is an auxiliary method that can be used in all the components. It takes the stream data and formats it to the given specifications.

It has different configuration states to configure its behaviour.

Expand Down Expand Up @@ -31,5 +31,3 @@ Code examples to show the behaviour of the **DataBuffer** class.
```python
--8<-- "docs/examples/others/data_buffer.py:example_3"
```

Go back [Index](../index.md)
20 changes: 10 additions & 10 deletions docs/auxiliar/persistency.md
Original file line number Diff line number Diff line change
Expand Up @@ -31,20 +31,20 @@ Two families ship today:
raw rows. Very storage heavy and *not recommended* for production-ready detectors.
- **Tracker backends** (`EventStabilityTracker`) keep only derived features
(e.g. "this variable has been constant for the last 10k events") that are relevant for the detector. Use these
when you only need a summary or a subset of the log's information, not the raw history -- they cost a fraction of
when you only need a summary or a subset of the log's information, not the raw history: they cost a fraction of
the memory.

All backends implement the same four-method contract: `add_data`, `get_data`,
`dump`, `load`. That contract is what `EventPersistency` and
`PersistencySaver` rely on -- anything you add later only has to follow it.
`PersistencySaver` rely on, so anything you add later only has to follow it.

### 3. Saver lifecycle (`PersistencySaver`)

`EventPersistency` itself is in-memory. To survive a process restart, the
state has to be written somewhere. `PersistencySaver` wraps an
`EventPersistency` and:

- writes to disk (or any `fsspec` URI) on two triggers -- a wall-clock interval
- writes to disk (or any `fsspec` URI) on two triggers: a wall-clock interval
and an event-count threshold;
- optionally `auto_load`s previously saved state during construction;
- exposes `start()` / `stop()` so the background timer can be torn down
Expand Down Expand Up @@ -106,7 +106,7 @@ ep[event_id] # alias for get_event_data
|---|---|
| `persistency.EventStabilityTracker` | You only care about how variables behave over time (`STATIC` / `STABLE` / `UNSTABLE` / `RANDOM`). Cheapest memory footprint. |

All three are re-exported from the top of the package -- `persistency.X` is the
All three are re-exported from the top of the package: `persistency.X` is the
canonical import; the deeply nested submodules are an implementation detail.

### Persisting to disk
Expand All @@ -129,7 +129,7 @@ saver.stop() # final flush, stops the background timer

`PersistencySaver.save()` is thread-safe, and `stop()` is idempotent. The two
save triggers (`save_interval_seconds` and `events_until_save`) are
independent -- whichever fires first wins.
independent: whichever fires first wins.

If writing to storage fails (unwritable path, full disk, lost credentials),
`save()` and `stop()` raise `persistency.PersistencySaveError`, and
Expand All @@ -149,21 +149,21 @@ saver = persistency.PersistencySaver(
```

If `auto_load=True` and no saved state exists, the constructor raises
`persistency.PersistencyLoadError` immediately -- fail-fast rather than
`persistency.PersistencyLoadError` immediately: it fails fast rather than
silently starting empty.

#### Exporting and importing state on demand

For one-shot transfers -- e.g. moving trained state to a new environment, or
taking a manual snapshot -- use the standalone functions directly:
For one-shot transfers (e.g. moving trained state to a new environment, or
taking a manual snapshot), use the standalone functions directly:

```python
from detectmatelibrary.utils import persistency

# Export to a file URI
persistency.save(ep, "./snapshots/trained-state")

# Export to bytes (no disk I/O -- useful when sending state over a network API)
# Export to bytes (no disk I/O; useful when sending state over a network API)
data: bytes = persistency.save(ep)

# Import from a file URI
Expand All @@ -187,7 +187,7 @@ when a saver is active.
#### Detector-level export and import

When working through a detector (the typical path for DetectMateService), use
the methods on the detector object directly -- no need to access
the methods on the detector object directly, with no need to access
`EventPersistency` internals:

```python
Expand Down
14 changes: 7 additions & 7 deletions docs/contribution.md
Original file line number Diff line number Diff line change
Expand Up @@ -12,7 +12,7 @@ at least the following information in a bug report:

1. Description of the bug. Describe the problem clearly.
2. Steps to reproduce. With the following configuration, go to.., click.., see error
3. Expected behavoir. What should happen?
3. Expected behavior. What should happen?
4. Environment. What was the environment for the test(version, browser, etc..)

*Please don't include any private/sensitive information in your issue! For reporting security-related issues, see [SECURITY.md](https://github.com/ait-detectmate/DetectMateLibrary/blob/main/SECURITY.md)*
Expand All @@ -39,7 +39,7 @@ git clone -b development git@github.com:YOURUSERNAME/DetectMateLibrary.git

### 3. Create a feature branch

Every single workpackage should be developed in it's own feature-branch. Use a name that describes the feature:
Every single workpackage should be developed in its own feature-branch. Use a name that describes the feature:

```bash
cd DetectMateLibrary
Expand All @@ -48,7 +48,7 @@ git checkout -b feature-some_important_work

### 4. Develop your feature and improvements in the feature-branch

Please make sure that you commit only improvements that are related to the workpage you created the feature-branch for.
Please make sure that you commit only improvements that are related to the workpackage you created the feature-branch for.

### 5. Fetch and merge from the upstream

Expand Down Expand Up @@ -103,23 +103,23 @@ git rebase -i HEAD~2

Delete your local feature-branch after the pull-request was merged into the development branch.

### 8. Update your local main branch
### 8. Update your local development branch

Update your local development branch:

```bash
git fetch upstream development
git checkout -b development
git checkout development
git rebase upstream/development
```

Additional infos:

- [https://www.atlassian.com/git/tutorials/merging-vs-rebasing](https://www.atlassian.com/git/tutorials/merging-vs-rebasing)

### 9. Update your main branch in your github-repository
### 9. Update the development branch in your GitHub repository

Please make sure that you updated your local development branch as described in section 8. above. After that push the changes to your github-repository to keep it up2date:
Please make sure that you updated your local development branch as described in section 8. above. After that push the changes to your github-repository to keep it up to date:

```bash
git push
Expand Down
Loading
Loading