Skip to content

Repository files navigation

Step 14: Replace the root README

Replace the existing root README.md with:

# Linux Observability Lab

[![Linux Observability Validation](https://github.com/solonsah/linux-observability-lab/actions/workflows/observability-validation.yml/badge.svg)](https://github.com/solonsah/linux-observability-lab/actions/workflows/observability-validation.yml)

A practical Linux observability lab for monitoring system health, performance, logs, and service availability.

This project demonstrates read-only system discovery, Prometheus metric collection, Grafana visualization, alert
design, incident triage, security controls, and automated configuration validation.

> This repository uses fictional or loopback targets and contains no employer, customer, credential, production
> log, or identifying infrastructure data.

## Project Objectives

- Collect Linux CPU, memory, load, uptime, filesystem, and systemd health data
- Produce a small set of Prometheus-compatible metrics with Bash
- Perform read-only observability checks with Ansible
- Configure Prometheus scrape jobs and alert rules
- Visualize Linux health through a reusable Grafana dashboard
- Document safe alert investigation and escalation
- Validate Bash, YAML, JSON, Prometheus, and Ansible content automatically

## Architecture

```mermaid
flowchart LR
    A[Linux host] --> B[Node Exporter]
    A --> C[Bash and Ansible checks]
    B --> D[Prometheus]
    D --> E[Alert rules]
    D --> F[Grafana]
    E --> G[Operator response]
    C --> G
```

See the detailed [architecture documentation](docs/architecture.md).

## Repository Structure

```text
.
├── .github/
│   └── workflows/
│       └── observability-validation.yml
├── ansible/
│   └── observability-precheck.yml
├── config/
│   └── prometheus/
│       ├── alerts.yml
│       └── prometheus.yml
├── dashboards/
│   └── grafana/
│       └── linux-overview.json
├── docs/
│   ├── alert-response.md
│   └── architecture.md
├── sample-output/
│   └── README.md
├── scripts/
│   └── collect-system-metrics.sh
├── .gitattributes
├── .gitignore
├── LICENSE
└── README.md
```

## Bash Metrics Collector

`scripts/collect-system-metrics.sh` reads Linux health information from:

- `/proc/loadavg`
- `/proc/uptime`
- `/proc/meminfo`
- `df`
- `systemctl`

It reports:

- Collection time
- One-minute load average
- System uptime
- Online logical CPU count
- Total and available memory
- Root-filesystem size and available space
- Number of failed systemd units

Run it on an authorized Linux lab system:

```bash
bash scripts/collect-system-metrics.sh
```

Use a non-sensitive label when needed:

```bash
HOST_LABEL="linux-lab" bash scripts/collect-system-metrics.sh
```

The script prints Prometheus text format to standard output. It is an educational collector and is not automatically
scraped by the included Prometheus configuration. Production monitoring should normally use a maintained exporter
such as Node Exporter.

## Ansible Observability Precheck

The Ansible playbook performs read-only checks against the `observability_targets` inventory group:

```bash
ansible-playbook -i inventory.ini ansible/observability-precheck.yml
```

It collects:

- Operating-system and kernel information
- Uptime and load
- Memory information
- Filesystem utilization
- Failed systemd units
- Presence of Node Exporter, Prometheus, and Grafana executables

The playbook does not install software, modify configuration, restart services, or use privilege escalation.

## Prometheus Configuration

`config/prometheus/prometheus.yml` provides a safe lab example with:

- A 15-second scrape interval
- A 15-second evaluation interval
- A 10-second scrape timeout
- Loopback targets for Prometheus and Node Exporter
- Separate `prometheus` and `linux_nodes` jobs
- External alert rules loaded from `alerts.yml`

The example assumes:

- Prometheus is available on `127.0.0.1:9090`
- Node Exporter is available on `127.0.0.1:9100`

These values are examples and must be reviewed before use outside a local lab.

## Alert Rules

| Alert | Condition | Duration | Severity |
|---|---|---:|---|
| LinuxNodeDown | Linux target is unreachable | 2 minutes | Critical |
| HighCpuUtilization | CPU utilization exceeds 85% | 10 minutes | Warning |
| HighMemoryUtilization | Memory utilization exceeds 90% | 10 minutes | Warning |
| LowRootFilesystemSpace | Root filesystem utilization exceeds 85% | 15 minutes | Warning |
| FailedSystemdUnit | A systemd unit remains failed | 5 minutes | Warning |

The thresholds are lab defaults. Production thresholds should be based on service-level objectives, workload
baselines, maintenance behavior, business impact, and alert-noise testing.

See the [alert-response runbook](docs/alert-response.md) for read-only diagnostic steps, interpretation,
escalation criteria, and post-change validation.

## Grafana Dashboard

`dashboards/grafana/linux-overview.json` provides panels for:

- Node availability
- CPU utilization
- Memory utilization
- Root-filesystem utilization
- Failed systemd units

To import it:

1. Open Grafana.
2. Go to **Dashboards**.
3. Select **New** and then **Import**.
4. Upload `dashboards/grafana/linux-overview.json`.
5. Select the appropriate Prometheus data source.
6. Import the dashboard.

The dashboard contains no real endpoints, credentials, or infrastructure identifiers.

The failed-systemd-units panel and related alert require the Node Exporter systemd collector to be enabled and
permitted to read the required systemd information.

## Metrics, Logs, and Alerts

Metrics show numerical behavior over time, such as CPU utilization or filesystem capacity.

Logs provide detailed events, such as service failures, authentication errors, kernel messages, and application
exceptions.

Alerts identify sustained conditions requiring investigation. An alert should be tied to impact, ownership,
duration, escalation, and a documented response.

This repository focuses primarily on metrics and read-only system investigation. It documents log analysis but does
not deploy a production log-aggregation platform.

## Automated Validation

GitHub Actions validates:

- Bash scripts with ShellCheck
- YAML formatting with yamllint
- Grafana dashboard JSON with `jq`
- Prometheus configuration and alert rules with `promtool`
- Ansible playbook syntax with `ansible-playbook --syntax-check`

The workflow validates repository content only. It does not connect to external infrastructure or deploy monitoring
services.

## Security and Privacy Controls

The repository excludes:

- Credentials, tokens, private keys, and environment files
- Private inventories and Ansible variables
- Prometheus and Grafana runtime databases
- Generated metrics, logs, reports, and artifacts
- Backups and deployment archives

Additional safeguards include:

- Loopback-only example targets
- Read-only scripts and playbooks
- No privilege escalation
- No embedded authentication material
- Sanitization requirements for all published output
- Least-privilege and restricted-network recommendations

Monitoring endpoints and administrative interfaces should not be publicly exposed without appropriate
authentication, authorization, TLS, network controls, and logging.

## Operational Safety

Before acting on an alert:

1. Confirm the environment and affected service.
2. Begin with read-only diagnostics.
3. Determine impact, cause, ownership, and urgency.
4. Obtain required approval before restarts, cleanup, configuration changes, or capacity changes.
5. Define rollback and validation requirements.
6. Validate both system health and application health after the change.
7. Document the cause, action, approvals, and outcome.

An alert returning to green does not prove that the application is healthy.

## Data Protection

Never publish raw workplace dashboards, logs, alerts, screenshots, reports, hostnames, domains, IP addresses,
usernames, application names, monitoring URLs, credentials, tokens, or identifying infrastructure details.

See the [sample-output guidance](sample-output/README.md).

## Technologies

- Linux
- Bash
- Ansible
- Prometheus
- PromQL
- Node Exporter
- Grafana
- YAML
- JSON
- ShellCheck
- yamllint
- GitHub Actions

## Project Limitations

This repository is a portfolio lab, not a complete production observability platform. Production environments also
require high availability, durable storage, retention planning, authentication, authorization, TLS, notification
routing, maintenance silences, backups, capacity management, and tested incident procedures.

## Project Status

Active portfolio project focused on safe, explainable, and repeatable Linux observability practices.

## License

This project is licensed under the [MIT License](LICENSE).

This README includes the explanations we developed and links to the detailed architecture, response runbook, and sanitization guidance. Save it with one clean newline at the end.

About

A practical Linux observability lab for monitoring system health, performance, logs, and service availability.

Topics

Resources

Stars

0 stars

Watchers

0 watching

Forks

Releases

Packages

Contributors

Languages