Skip to content
Merged
Show file tree
Hide file tree
Changes from all commits
Commits
File filter

Filter by extension

Filter by extension


Conversations
Failed to load comments.
Loading
Jump to
Jump to file
Failed to load files.
Loading
Diff view
Diff view
2 changes: 2 additions & 0 deletions .github/workflows/trufflehog-evidence-example.yml
Original file line number Diff line number Diff line change
Expand Up @@ -55,6 +55,8 @@ jobs:
run: |
ls
python ./examples/trufflehog/jsonl_to_json_converted.py
# The JSONL still holds Raw/RawV2. Only the sanitized predicate and markdown are published.
rm -f trufflehog-results.jsonl
jf evd create \
--package-name $IMAGE_NAME \
--package-version $VERSION \
Expand Down
26 changes: 17 additions & 9 deletions examples/trufflehog/README.md
Original file line number Diff line number Diff line change
Expand Up @@ -4,6 +4,12 @@ This repository provides a working example of a GitHub Actions workflow that aut

This workflow is an essential DevSecOps practice, helping to prevent accidental secret leakage by creating a traceable and auditable record of what was found in your codebase at a specific point in time.

### **Do not publish secret material as evidence**

Evidence created with `jf evd create` is stored on the package. Anyone who can read that package or repository can read the evidence, including other teams, service accounts, and anonymous readers when the repository allows it. That audience is wider than the people who can read the source repository. Signed evidence also cannot be quietly edited later, so a secret written into it stays there.

TruffleHog's JSON output includes the matched secret in `Raw` and `RawV2`, and sometimes again in `ExtraData`. This example strips those fields before anything is attached. The predicate and the optional Markdown report keep only non-secret metadata: detector name, source location, the `Redacted` value, and verification status.

### **Key Features**

* **Automated Secret Scanning**: Uses Trufflehog to scan the repository for potential secrets and sensitive information.
Expand Down Expand Up @@ -98,8 +104,8 @@ Once the workflow completes successfully, you can navigate to your repository in

1. **Setup and Checkout**: The workflow begins by setting up the JFrog CLI and checking out the repository code using sparse checkout to focus on the trufflehog example directory.
2. **Run Trufflehog Secret Scan**: Uses Docker to run Trufflehog against the repository, scanning for potential secrets and sensitive information. The scan outputs results in JSON format.
3. **Process Scan Results**: A Python helper script (`process_trufflehog_results.py`) parses the Trufflehog JSON output and generates a structured predicate file suitable for JFrog Evidence.
4. **Generate Optional Markdown Report**: If ATTACH_OPTIONAL_CUSTOM_MARKDOWN_TO_EVIDENCE is true, the Python script creates a human-readable Markdown report summarizing the findings.
3. **Build the evidence predicate**: `jsonl_to_json_converted.py` turns the JSONL scan into `trufflehog.json` and drops secret fields (`Raw`, `RawV2`, `ExtraData`, and anything outside the allowlist) before the file is attached.
4. **Generate Optional Markdown Report**: If ATTACH_OPTIONAL_CUSTOM_MARKDOWN_TO_EVIDENCE is true, `process_trufflehog_results.py` writes a human-readable summary that includes the redacted value and location, and does not include the raw secret.
5. **Attach Signed Evidence**: The final step uses the jf evd create command to attach the scan results as evidence to the specific package version in Artifactory. The evidence is signed using the provided private key, ensuring its authenticity and integrity.

### **Key Commands Used**
Expand All @@ -108,17 +114,18 @@ Once the workflow completes successfully, you can navigate to your repository in
This step runs the `trufflesecurity/trufflehog` container to scan the entire checked-out repository. The results are output in a `.jsonl` (JSON Lines) format. The `|| true` ensures the workflow continues even if secrets are found, allowing the findings to be reported as evidence.

```bash
docker run --rm -it -v "$PWD:/pwd" trufflesecurity/trufflehog:latest filesystem /pwd --json
docker run --rm -v "$PWD:/pwd" trufflesecurity/trufflehog:latest filesystem /pwd --json > trufflehog-results.jsonl
```

* **Process Results:**
The raw `.jsonl` output from Trufflehog is processed in two steps:
The raw `.jsonl` output from Trufflehog is processed in two steps. Both scripts remove secret fields before writing files that `jf evd create` will publish.

1. A Python script (`jsonl_to_json_converted.py`) converts the JSON Lines file into a standard, well-formed JSON array named `trufflehog.json`, which is required for the evidence predicate.
2. If `ATTACH_OPTIONAL_CUSTOM_MARKDOWN_TO_EVIDENCE` is `true`, a second script (`process_trufflehog_results.py`) generates a human-readable Markdown summary.
1. `jsonl_to_json_converted.py` converts the JSON Lines file into `trufflehog.json`, the evidence predicate. Each finding keeps only the allowlisted non-secret fields.
2. If `ATTACH_OPTIONAL_CUSTOM_MARKDOWN_TO_EVIDENCE` is `true`, `process_trufflehog_results.py` generates a human-readable Markdown summary without `Raw` or `RawV2`.

```bash
python process_trufflehog_results.py trufflehog-results.json
python jsonl_to_json_converted.py trufflehog-results.jsonl trufflehog.json
python process_trufflehog_results.py trufflehog-results.jsonl
```

* **Attach Evidence:**
Expand All @@ -131,8 +138,9 @@ jf evd create \
--package-repo-name your-repo-name \
--key "${{ secrets.JF_PRIVATE_KEY }}" \
--key-alias ${{ vars.JF_SIGNING_KEY_ALIAS }} \
--predicate ./trufflehog-evidence.json \
--predicate-type http://trufflesecurity.com/trufflehog/secret-scan
--predicate ./trufflehog.json \
--predicate-type https://trufflesecurity.com/TruffleHog \
--markdown report_readme.md
```

### **References**
Expand Down
54 changes: 33 additions & 21 deletions examples/trufflehog/jsonl_to_json_converted.py
Original file line number Diff line number Diff line change
@@ -1,23 +1,35 @@
import json
import sys

# Input and output file paths
input_file = "trufflehog-results.jsonl"
output_file = "trufflehog.json"

# Read the JSONL file and convert it to a list of JSON objects
with open(input_file, "r") as infile:
data = [json.loads(line) for line in infile]

# Wrap the data in a dictionary with the key "data"
output_data = {"data": data}

# Write the output to a JSON file with proper formatting
with open(output_file, "w") as outfile:
outfile.write('{\n "data": [\n')
for i, item in enumerate(data):
json.dump(item, outfile, indent=4)
if i < len(data) - 1:
outfile.write(',\n')
outfile.write('\n ]\n}')

print(f"Converted {input_file} to {output_file}")
from sanitize_trufflehog import sanitize_record

DEFAULT_INPUT_FILE = "trufflehog-results.jsonl"
DEFAULT_OUTPUT_FILE = "trufflehog.json"


def convert(input_file, output_file):
"""Write an evidence predicate that keeps only non-secret finding fields."""
records = []
with open(input_file, "r", encoding="utf-8") as infile:
for line in infile:
stripped = line.strip()
if not stripped:
continue
records.append(sanitize_record(json.loads(stripped)))

with open(output_file, "w", encoding="utf-8") as outfile:
json.dump({"data": records}, outfile, indent=2)
outfile.write("\n")

print(f"Converted {input_file} to {output_file}")
return records


def main(argv):
input_file = argv[1] if len(argv) > 1 else DEFAULT_INPUT_FILE
output_file = argv[2] if len(argv) > 2 else DEFAULT_OUTPUT_FILE
convert(input_file, output_file)


if __name__ == "__main__":
main(sys.argv)
64 changes: 26 additions & 38 deletions examples/trufflehog/process_trufflehog_results.py
Original file line number Diff line number Diff line change
@@ -1,25 +1,23 @@
import json
import sys

from sanitize_trufflehog import sanitize_record, source_location


def generate_markdown_report(report):
source_name = report.get('SourceName', 'N/A')
detector_name = report.get('DetectorName', None)
"""Build a Markdown finding that contains no raw secret material."""
sanitized = sanitize_record(report)
detector_name = sanitized.get("DetectorName")
if not detector_name:
return None # Skip entries without 'DetectorName'

detector_description = report.get('DetectorDescription', 'N/A')
verified = report.get('Verified', False)
raw = report.get('Raw', 'N/A')
redacted = report.get('Redacted', 'N/A')
return None

extra_data = report.get('ExtraData', {})
account = extra_data.get('account', 'N/A')
arn = extra_data.get('arn', 'N/A')
is_canary = extra_data.get('is_canary', 'N/A')
message = extra_data.get('message', 'N/A')
resource_type = extra_data.get('resource_type', 'N/A')
source_name = sanitized.get("SourceName", "N/A")
detector_description = sanitized.get("DetectorDescription", "N/A")
verified = sanitized.get("Verified", False)
redacted = sanitized.get("Redacted", "N/A")
location = source_location(sanitized)

markdown_report = f"""
return f"""
## Report Overview: {source_name}

**Source Name:** `{source_name}`
Expand All @@ -28,50 +26,40 @@ def generate_markdown_report(report):

**Detector Description:** `{detector_description}`

**Verified:** `{verified}`
**Source Location:** `{location}`

**Raw Data:** `{raw}`
**Verified:** `{verified}`

**Redacted Data:** `{redacted}`

---
### Extra Data
| Key | Value |
| :------------ | :------------ |
| Account | {account} |
| ARN | {arn} |
| Is Canary | {is_canary} |
| Message | {message} |
| Resource Type | {resource_type} |
---
"""
return markdown_report


def main(input_file):
# Read JSONL input from a file
with open(input_file, 'r') as file:
with open(input_file, "r", encoding="utf-8") as file:
lines = file.readlines()

markdown_reports = []
for line in lines:
report = json.loads(line)
stripped = line.strip()
if not stripped:
continue
report = json.loads(stripped)
markdown_report = generate_markdown_report(report)
if markdown_report:
markdown_reports.append(markdown_report)

# Define the output file path
output_file = 'report_readme.md'

# Write the Markdown reports to a file
with open(output_file, 'w') as file:
output_file = "report_readme.md"
with open(output_file, "w", encoding="utf-8") as file:
file.write("\n\n".join(markdown_reports))

print(f"Markdown README generated successfully and saved to {output_file}!")

if __name__ == '__main__':

if __name__ == "__main__":
if len(sys.argv) != 2:
print("Usage: python process_trufflehog_results.py <report_file>")
sys.exit(1)

input_file = sys.argv[1]
main(input_file)
main(sys.argv[1])
100 changes: 100 additions & 0 deletions examples/trufflehog/sanitize_trufflehog.py
Original file line number Diff line number Diff line change
@@ -0,0 +1,100 @@
"""Strip secret material from TruffleHog findings before they become evidence.

Evidence is readable by every principal who can read the subject package, and a
signed predicate cannot be edited in place. TruffleHog puts the matched secret
in Raw/RawV2 and often again in ExtraData, so the published record is an
allowlist of non-secret fields.
"""

# Fields safe to publish. Anything else (Raw, RawV2, ExtraData, StructuredData,
# VerificationError) is dropped, including fields added by newer TruffleHog versions.
SAFE_FIELDS = (
"DetectorName",
"DetectorType",
"DetectorDescription",
"DecoderName",
"Verified",
"SourceName",
"SourceType",
"SourceID",
"SourceMetadata",
"Redacted",
)

_SECRET_VALUE_KEYS = ("Raw", "RawV2")
# Replacing a very short secret as a substring would corrupt unrelated text.
_MIN_SUBSTRING_SECRET_LENGTH = 8
_REDACTION_PLACEHOLDER = "[redacted]"


def sanitize_record(record):
"""Return a copy of a TruffleHog finding with secret fields removed."""
if not isinstance(record, dict):
raise TypeError("TruffleHog finding must be a JSON object")

secrets = _secret_values(record)
sanitized = {}
for field in SAFE_FIELDS:
if field not in record:
continue
sanitized[field] = _scrub(record[field], secrets)
return sanitized


def source_location(record):
"""Human-readable file:line (or equivalent) from SourceMetadata."""
metadata = record.get("SourceMetadata") if isinstance(record, dict) else None
data = metadata.get("Data") if isinstance(metadata, dict) else None
if not isinstance(data, dict) or not data:
return "N/A"

locations = []
for details in data.values():
if not isinstance(details, dict):
continue
path = details.get("file") or details.get("path") or ""
line = details.get("line")
if path and line is not None:
locations.append(f"{path}:{line}")
elif path:
locations.append(str(path))
elif line is not None:
locations.append(f"line {line}")
return ", ".join(locations) if locations else "N/A"


def _secret_values(record):
secrets = []
seen = set()
for key in _SECRET_VALUE_KEYS:
value = record.get(key)
if isinstance(value, str) and value and value not in seen:
seen.add(value)
secrets.append(value)
secrets.sort(key=len, reverse=True)
return secrets


def _scrub(value, secrets):
if isinstance(value, str):
return _scrub_string(value, secrets)
if isinstance(value, list):
return [_scrub(item, secrets) for item in value]
if isinstance(value, dict):
cleaned = {}
for key, item in value.items():
if key in _SECRET_VALUE_KEYS:
continue
cleaned[key] = _scrub(item, secrets)
return cleaned
return value


def _scrub_string(value, secrets):
redacted = value
for secret in secrets:
if redacted == secret or (
len(secret) >= _MIN_SUBSTRING_SECRET_LENGTH and secret in redacted
):
redacted = redacted.replace(secret, _REDACTION_PLACEHOLDER)
return redacted
Loading
Loading