Skip to content
CyberSmithSECURE
Under Attack

Configuration Review

Configuration Review of Firewalls

A firewall rule base is a historical document. Every rule was added for a reason, most of those reasons have expired, and almost nobody removes anything. The review therefore treats the rule base as an archaeology problem — what does this actually permit today, versus what does the organisation believe it permits — and the gap between those two is the finding.

Methodology

  1. 01

    Configuration collection

    Running configuration exported from every device in scope, along with object groups, NAT rules, routing and the management plane settings.

  2. 02

    Platform hardening review

    The device itself assessed against CIS benchmark and vendor guidance: management access, authentication, logging, time synchronisation, firmware currency and high-availability configuration.

  3. 03

    Rule base analysis

    Every rule evaluated for shadowing, redundancy, over-permissive source, destination or service, and absence of logging. Rule order is analysed because a shadowed rule is a rule that does nothing.

  4. 04

    Policy intent reconciliation

    The effective permitted flows computed and compared against the documented network policy. Where no documented policy exists, that is the primary finding.

  5. 05

    Rule usage analysis

    Hit counts reviewed over an agreed period to identify rules that have never matched, which are candidates for removal with evidence rather than guesswork.

  6. 06

    Change and approval review

    How rules get added, who approves them, whether there is an expiry mechanism, and whether emergency changes are reconciled afterwards.

  7. 07

    Reporting and retest

    Technical report and executive summary together, with a prioritised rule remediation list, then a retest confirming closure.

Approach to testing

  • Offline analysis of exported configuration by default. Nothing is changed and nothing is sent to the device, so the review carries no outage risk.
  • Rules are assessed against intent, not only against a benchmark. A rule permitting any-any between two internal zones is a finding regardless of what CIS says about the platform.
  • Hit count data is requested for at least ninety days. Recommending removal of a rule that fires quarterly is how a reviewer causes an outage from a desk.
  • Every recommended change carries a stated blast radius, because a rule base change on a perimeter device is a production change.
  • Where the client runs several vendors, the review normalises findings so they are comparable rather than reported in three vendor dialects.

Types of assessment

Rule base review (default)

Offline analysis of exported configuration. No device interaction, no risk, highest finding yield per unit of effort.

Platform hardening review

The device itself rather than its policy: management plane, authentication, logging and firmware. Often bundled with the rule base review.

Segmentation validation

Rule base review paired with active testing that confirms the rules behave as written. Frequently required for PCI DSS scope reduction.

Continuous policy monitoring

Configuration exported and re-analysed on a schedule, so drift and emergency changes are caught rather than discovered annually.

Frameworks and standards

CIS Benchmarks
Platform-specific hardening baselines for Palo Alto, Fortinet, Cisco ASA and Check Point.
Vendor hardening guides
Manufacturer guidance, which is often stricter and more current than the benchmark.
NIST SP 800-41
Guidelines on firewalls and firewall policy, used for rule base structure and policy design.
PCI DSS Requirement 1
Where the client is in scope, rule review and the six-month review requirement are assessed directly.
ISO/IEC 27001 Annex A
Control mapping where the client maintains an ISMS.

Tools used

Tooling is where testing starts, not where it ends. Every automated result is reproduced by hand before it reaches a report.

Nipper

Automated device configuration audit against CIS and vendor baselines.

Vendor management platforms

Panorama, FortiManager and similar, read-only, for rule usage and hit count data.

Custom rule base parsers

Normalising configuration from several vendors into one comparable model.

Nmap

Confirming that the rule base behaves as written, where segmentation validation is in scope.

Algosec / Tufin (where deployed)

Reading the client's existing policy analysis rather than duplicating it.

Checklist approach

The checklist is the floor, not the ceiling. It guarantees coverage so nothing standard is missed; the findings that matter usually come from what a tester does after it is complete.

Rule base hygiene

  • Any-any rules on source, destination or service
  • Shadowed rules that can never match
  • Redundant and duplicate rules
  • Rules with no logging enabled
  • Rules with no owner, ticket reference or expiry
  • Zero-hit rules over the review period

Perimeter policy

  • Inbound rules against documented business need
  • Management protocols permitted from untrusted zones
  • Outbound egress filtering, or its absence
  • Permitted protocols against a current threat baseline
  • Geo and reputation filtering where licensed

Internal segmentation

  • Zone-to-zone permitted flows against the design
  • Server-to-server rules that should be host-based
  • User network access to management networks
  • Guest and BYOD containment
  • OT and corporate separation where both exist

Platform hardening

  • Management interface exposure and access control
  • Administrative authentication, MFA and role separation
  • Firmware version against vendor advisories
  • Logging destination, retention and integrity
  • Time synchronisation and its effect on log correlation
  • High availability and failover configuration

Governance

  • Change approval process and its evidence
  • Rule expiry and periodic review mechanism
  • Emergency change reconciliation
  • Configuration backup and restoration testing
  • Documented network policy the rule base can be assessed against

How findings are scored

Every finding is scored on CVSS 3.1 and placed in one of five levels. The executive summary adds a sixth band — Compliant — so components that passed appear on the same chart as those that did not.

Critical
Immediate measures must be taken. These vulnerabilities can allow an attacker to take complete control of the application or server — stealing user data, tricking users into supplying sensitive information, or defacing the site.
High
Maximum risk associated with a specific vulnerability instance. May enable an attacker to compromise the application and its data, partially or completely, or to modify application behaviour beyond its intended purpose. To be handled with utmost priority.
Medium
Considerable risk. May enable an attacker to exploit the application to a particular level, gaining low-level information that can be used to craft more specific attacks.
Low
Lowest risk. May allow an attacker to gain some information about the application that was not intended to be known, without an exploitation technique currently available at that instance.
Informational
A functionality or component is missing best-practice implementation. Not a risk today, but may become one as the application changes or as exploitation techniques, policy or legal requirements evolve.

Scan types selected

  • Safe Checks
  • Standard / OWASP Top 10
  • Destructive
  • SANS Top 25
  • Business Logic Vulnerability Testing

Standard toolset by stage

OSINT
Datasploit, Google Dorks, Shodan
Enumeration & Scanning
Nmap, Wfuzz, Unicornscan
Domain Enumeration
Nikto, DnsRecon, Knock
Crawling & Fuzzing
Burp Suite, Acunetix, Netsparker
Vulnerability Analysis
OpenSSL, sqlmap, CVE-Details
Exploitation
Metasploit, Netcat, Exploit-DB

How CSS tests

A unified swarm of agents, for blind spot detection

AI agents drive several testing tracks against the same target at once, then cross-check each other. A single tester works one hypothesis at a time; parallel agents cover the space a sequential pass leaves behind.

  • Shadowing analysis is a pairwise comparison across the entire rule base — a thousand rules is half a million comparisons, which a human reviewer samples and a machine completes.

  • Object group expansion hides scope. A rule that looks narrow can permit a /16 once nested groups are resolved, and resolving every group across every rule is mechanical work.

  • Multi-vendor estates need normalisation before comparison, otherwise each device is reviewed in isolation and the path that crosses two vendors is never seen.

  • Hit count correlation against rule intent identifies removal candidates with evidence, rather than a reviewer's impression that a rule looks unused.

Every agent finding is validated by a human reviewer, and no recommendation reaches the report without a stated blast radius. The swarm analyses configuration; it never connects to a device.

Why this differs

What CSS does that most vendors do not

Every one of these is checkable. Ask any vendor for the same and compare the answers.

Intent, not just benchmark

A device can pass every CIS check and still permit any-any between the office and the card data environment. The review assesses what the rule base actually allows against what the organisation intends.

Evidence for removal

Rules are recommended for removal with hit count evidence over an agreed window, not on the reviewer's impression. That is the difference between a report that gets actioned and one that gets shelved.

Both reports, always

Technical report and executive summary together, plus a prioritised change list the network team can work from directly.

Fixation, not a backlog

CSS works the change sequence with the network team, including which changes need a maintenance window, and retests to evidence closure.

Reporting

Two documents, two audiences

Both are produced for every engagement. They are not the same document at two lengths — they answer different questions and are written separately. The structure below is the one CSS actually issues.

Technical assessment report

For the engineers who will fix it

  • Disclaimer, and Limitations on Disclosure and Use
  • Risk Level & Description — the five levels above, scored on CVSS 3.1
  • Scan Type — which of the five assessment types were selected
  • Assessment Scope — the control areas covered
  • Assessment Date — the exact testing window
  • Objective of the Assessment — objectives listed against completion status
  • Tools Utilization — manual and automated tooling by stage
  • Summary of the Assessment
  • Overall Recommendations, split into Must Have and Should Have
  • Vulnerability Overall Classifications as per Organization
  • Security Issues Highlighted
  • The Key Findings — each with evidence and detailed recommendation
  • Summary of Findings & Conclusion

For this assessment specifically

  • Scope: devices, firmware versions, configuration export date
  • Findings against CIS and vendor baseline, with benchmark references
  • Full rule base analysis: shadowed, redundant, over-permissive and unlogged rules, listed by rule number
  • Effective permitted flows between zones, compared with documented policy
  • Zero-hit rules with the observation window stated
  • Prioritised change list with blast radius for each change
  • Retest results appended against each original finding

Executive summary

For the people who will fund the fix

  • Objectives, each against a completion status
  • Overall Finding of the Assessment — total threats identified, broken down by component and severity
  • Summary of the Assessment
  • Artefacts of the Assessment — the key findings as a numbered register with severity
  • Observation of the Assessment — the major attacks the organisation should be prepared for, given what was found
  • Overall Recommendation, including a Business Enabling Recommendation sequence
  • Must Have and Should Have actions

For this assessment specifically

  • What the firewall actually permits, against what the organisation believes it permits
  • Number of rules, and how many are shadowed, redundant or unused
  • The three changes that most reduce exposure
  • Compliance position where PCI DSS or an ISMS applies
  • Remediation timeline and retest date
  • One page

Case studies

What this finds in practice

Representative engagement patterns. Sector and scale only — no client is named, and no detail is included that could identify one.

A retail group with 140 stores behind a central perimeter pair.

Finding
The rule base had grown to 1,840 rules over eleven years. Analysis found 312 shadowed rules that could never match, 96 with no logging, and an any-any rule between the store network and the card data environment added during a 2019 migration and never removed.
Recommendation
Remove the any-any rule immediately with a narrow replacement, then run a staged clean-up of shadowed and unlogged rules against hit count evidence.
Outcome
The any-any rule was replaced within 24 hours. Clean-up removed 430 rules over three months with no service impact, and the base is now reviewed quarterly.

A manufacturer running Fortinet at the perimeter and Cisco ASA internally.

Finding
Each device passed its own vendor review. Normalising both rule bases into one model showed that a permitted flow through the perimeter, combined with an internal rule, allowed a supplier VPN to reach the plant historian — a path neither review saw in isolation.
Recommendation
Assess the estate as one policy rather than device by device, and add an explicit deny between supplier VPN and any OT-adjacent zone.
Outcome
Deny rule added at the next change window. Reviews are now scoped across the estate rather than per device.

A financial services firm preparing for a PCI DSS assessment.

Finding
Management interfaces on both perimeter firewalls accepted connections from the general user VLAN, and administrative accounts were local with no MFA. The six-monthly rule review required by PCI had been recorded as complete but produced no evidence.
Recommendation
Restrict management to a dedicated administrative network, federate administrative authentication with MFA, and produce a documented review with rule dispositions as the evidence artefact.
Outcome
All three completed before the assessment. The documented review became the template for subsequent cycles, which is what made the control sustainable rather than a one-off.

Next

Scope this assessment

Most scopes are settled in one call. Tell us what the application does and who uses it, and we will tell you what testing it properly involves.