Configuration

Configuration examples

Worked configurations you can copy, from a first collection to filtering, masking and fan-out.

Each example below is complete: paste it into C:\ProgramData\DPLens\config\agent.yaml, change the addresses and paths to yours, and restart the service. Every setting they use is described in the Configuration reference.

Check a file before you install it:

"C:\Program Files\DPLens\dplens.exe" --validate-only --config C:\path\to\agent.yaml

The simplest thing that works

One Windows Event Log channel to one collector, as newline-delimited JSON.

utc_timestamps: true

sources:
  - type: wel
    name: app-events
    channels: [Application]
    params:
      read_existing: "false"

destinations:
  - type: tcp
    name: collector
    params:
      address: "collector.example.com:2514"
      format: "ndjson"

pipelines:
  - name: default
    source: app-events
    destinations: [collector]

read_existing: "false" means new events only. Set it to "true" to read what is already in the channel first — useful once, when you are backfilling, and rarely what you want afterwards.

Syslog over TLS, with failover

The usual shape for a production SIEM: RFC 5424 syslog over TLS, a second collector to fall back to, and a gigabyte of cache to ride out an outage.

utc_timestamps: true

sources:
  - type: wel
    name: security-log
    channels: [Security]
    params:
      read_existing: "false"

destinations:
  - type: tls
    name: siem
    params:
      address: "siem.example.com:6514"
      failover_addresses: "siem-2.example.com:6514"
      sni: "siem.example.com"
      ca_file: "C:\\ProgramData\\DPLens\\ca\\siem-ca.pem"
      format: "syslog-5424"
      framing: "octet-counted"
      facility: "13"
      severity: "5"
      app_name: "dplens"
      cache_max_bytes: "1073741824"
      cache_full_policy: "block-upstream"

pipelines:
  - name: security-to-siem
    source: security-log
    destinations: [siem]

Certificate validation is on. Omit ca_file to use the Windows trust store, or point it at your own authority's bundle. Failover moves to the second address when the first stops accepting and moves back when it recovers.

Collecting several channels at once

One source can read a list of channels, which is simpler than one source per channel and keeps them in step.

sources:
  - type: wel
    name: windows-core
    channels:
      - Security
      - System
      - Application
      - Microsoft-Windows-Sysmon/Operational
      - Microsoft-Windows-PowerShell/Operational
    params:
      read_existing: "false"

Keeping only what you care about

The cheapest event is the one you never send. A filter stage keeps or drops events by a condition.

pipelines:
  - name: security-to-siem
    source: security-log
    destinations: [siem]
    stages:
      - type: filter
        name: keep-interesting
        rules:
          action: keep
          rule:
            all:
              - pred:
                  field: EventID
                  op: in
                  values: [4624, 4625, 4634, 4688, 4720]
              - not:
                  pred: { field: TargetUserName, op: regex, value: "\\$$" }

That keeps a short list of event IDs, and of those drops anything whose target user name ends in $ — machine accounts.

Conditions nest. all requires every child to match, any requires one, not inverts, and pred is a single test. The tests available are eq, ne, contains, regex, gt, ge, lt, le, in and cidr.

Collapsing repeats

When something fails over and over, you usually want to know it happened and how often — not to carry every copy.

stages:
  - type: filter
    name: failed-logons-only
    rules:
      action: keep
      rule:
        pred: { field: EventID, op: eq, value: 4625 }
  - type: aggregate
    name: dedup-logon-failures
    aggregate:
      group_by: [TargetUserName, IpAddress]
      window_secs: 60
      representative: first
      max_groups: 10000

Failures for the same user from the same address inside a 60-second window become one event carrying the count. max_groups bounds how many distinct combinations are tracked at once; beyond it, new groups pass through and are counted so the number is visible rather than silently lost.

Aggregation goes after parsing and before enrichment and masking, so the later stages run once per representative instead of once per duplicate.

Masking sensitive data

Masking runs before anything is written to disk or sent, so a value a rule rewrites is not written to the disk cache and is not sent.

stages:
  - type: mask
    name: redact-pii
    rules:
      - detector: email
        fields: ["*"]
        action: redact
      - detector: luhn
        fields: ["*"]
        action: partial
        keep_last: 4
      - detector: uk-ni
        fields: ["*"]
        action: hmac
        salt: "secret://mask/salt"
      - detector: regex
        pattern: "AKIA[0-9A-Z]{16}"
        fields: [message]
        action: token
        token: "[aws-key]"

Every rule needs fields, which says where to look:

fieldsWhere the rule looks
["*"]Every text field on the event
[message, user]Only the fields you name
[raw]The original record, before parsing

The built-in detectors are luhn (card numbers), email, ipv4, ipv6, us-ssn, uk-ni and phone. Use regex with a pattern of your own for anything else. User patterns run on an engine with no backtracking, so a pattern cannot be written that stalls the pipeline.

actionWhat you get
redactThe value is replaced entirely.
partialAll but the last keep_last characters are replaced — enough to recognise a card without carrying it.
tokenReplaced with a fixed marker you choose.
hmacReplaced with a keyed hash: the same input always gives the same output, so you can still correlate, and the value itself is not carried. Treat hashed values of low-entropy data — national identifiers, card numbers, phone numbers — as pseudonymised rather than anonymised.

hmac needs a salt, given as a secret handle. Seed the value once:

"C:\Program Files\DPLens\dplens.exe" --set-secret secret://mask/salt

Every masking decision is counted and written to the audit log, so you can evidence what was redacted and when.

Capping how much you send

stages:
  - type: rate-control
    name: cap-egress
    rate_control:
      eps: 2000
      burst: 4000
      buffer_max_events: 50000
      high_watermark: 0.9
      low_watermark: 0.5
      policy: drop_newest

Delivery is capped at 2,000 events a second, allowing short bursts of 4,000. Overflow is smoothed through a bounded buffer; when that reaches 90% full, shedding starts, and it stops again at 50%. Everything shed is counted.

This protects a licence or an indexer from one machine that has started shouting. Use it with care: it is the one stage that deliberately discards events, and the Overview page will show you exactly how many.

Normalising field names

If your SIEM expects a common schema, normalise at the agent rather than at search time.

stages:
  - type: normalise
    name: cim
    mapping:
      schema: cim
      map: windows-security
      validate: true
      overrides:
        - set: { field: index, value: wineventlog }

schema is cim or ocsf. overrides lets you set, copy, rename, cast, default or drop individual fields afterwards.

Tailing a log file

sources:
  - type: file-tail
    name: app-log
    params:
      path: "C:/logs/app*.log"
      recursive: "false"
      encoding: "auto"
      read_existing: "true"
      flush_partial_after_ms: "3000"

path takes a wildcard. Files that appear later are picked up; files that are rotated are followed. A partial last line is held briefly in case the rest is still being written, then emitted anyway so a stalled writer cannot hide an event indefinitely.

Watching files for change

sources:
  - type: fim
    name: critical-files
    params:
      source_tag: "fim"
    watch_items:
      - id: hosts-file
        path: 'C:\Windows\System32\drivers\etc\hosts'
        period_secs: 300
      - id: web-content
        path: 'C:\inetpub\wwwroot\**\*'
        recursive: true
        exclude: ["*\\temp\\*", "*.tmp"]
        period_secs: 3600
        size_threshold_bytes: 104857600

Each watch item is checked on its own schedule and reports what changed. Files larger than size_threshold_bytes are tracked by their properties — size, timestamps, permissions — rather than by reading them, so one enormous file cannot dominate a scan.

Use single quotes around Windows paths so backslashes stay literal.

Receiving syslog from other devices

sources:
  - type: syslog
    name: firewall-syslog
    params:
      listen: "0.0.0.0:514"
      protocol: "udp"
      allow_ips: "192.0.2.0/24,198.51.100.0/24"
      max_message_bytes: "65536"

Always set allow_ips on a receiver. Left empty it accepts from anywhere that can reach the port.

Sending to two places at once

destinations:
  - type: tls
    name: siem
    params:
      address: "siem.example.com:6514"
      format: "syslog-5424"
  - type: tcp
    name: archive
    params:
      address: "archive.example.com:2514"
      format: "ndjson"
      cache_max_bytes: "2147483648"

pipelines:
  - name: security-fanout
    source: security-log
    destinations: [siem, archive]
    backpressure: spill

Both destinations receive the same processed events. Each has its own queue, so one being down does not hold up the other.

Running two pipelines

pipelines:
  - name: security-to-siem
    source: security-log
    destinations: [siem]
    stages:
      - type: mask
        name: redact-pii
        rules:
          - detector: email
            fields: ["*"]
            action: redact

  - name: files-to-archive
    source: app-log
    destinations: [archive]

  - name: system-to-siem
    source: system-log
    destinations: [siem]
    enabled: false

Every enabled pipeline runs at the same time. The third is switched off: it stays in the file, is validated like any other, and starts the moment you enable it.

A pipeline that pauses instead of queueing

Sometimes you would rather stop collecting than build a backlog — on a machine with little disk, for example. Windows keeps the events in its own log until DPLens catches up.

pipelines:
  - name: system-to-siem
    source: system-log
    destinations: [siem]
    backpressure: pause_source

This only works where the destination is not shared with another pipeline using a different policy, and not on a destination that waits for acknowledgement.

Next