Every engagement starts the same way, whether the customer asks for a migration, a cost review or managed services: a structured look at the platform as it is. Not because we expect it to be broken, but because the answer to almost every follow-up question depends on it. The list below is what we work through. It is not complete, and it does not need to be. It is the set of checks that, in our experience, explains most of the pain a team is feeling.

Topology and versions

  • Version and support window. Which Splunk version runs where, and how far it is from end of support. Mixed versions across a cluster are a finding on their own.
  • Cluster layout. Search head cluster or standalone, indexer cluster or not, replication and search factors, multisite or single site. We draw it, because the diagram the team has is usually two years old.
  • Deployment server and app distribution. Which server classes exist, which apps go where, and whether anything is still deployed by hand.
  • Forwarder inventory. How many universal and heavy forwarders, which versions, and how many have not phoned home in the last week.

Indexing and storage

  • Licence usage per index and sourcetype, last 30 days. The top ten sourcetypes usually account for more than half the volume, and one or two of them are often noise nobody reads.
  • Index configuration. Retention, hot/warm/cold paths, frozen handling, SmartStore or not. We look for indexes with the default settings, because defaults were never chosen.
  • Bucket health. Small buckets, excessive bucket counts, and replication or search factor not met.
  • Disk headroom and IOPS. Indexers close to their thresholds explain a lot of intermittent search slowness.

Search and knowledge objects

  • Scheduled search load. Skipped searches, searches with long runtimes, and the concurrency picture over a day. Skipped searches mean alerts that never fired.
  • Expensive searches. Real-time searches, wildcard sourcetypes, searches over all time. These are usually a handful of dashboards and a few well-meant alerts.
  • Knowledge object sprawl. Orphaned objects, duplicate lookups, macros nobody remembers, and permissions that make everything global.
  • Data models and acceleration. Which are accelerated, whether acceleration completes, and whether CIM compliance is real or assumed.

Data onboarding

  • Timestamp and line-breaking problems. A search for events with future timestamps or timestamps at midnight finds them fast.
  • Sourcetype hygiene. Sourcetypes ending in `-too_small`, `-N` suffixes, and inputs without explicit sourcetype or index.
  • Props and transforms. Whether parsing happens where it should (heavy forwarder or indexer), and whether index-time field extractions are justified.

Security and operations

  • Authentication and roles. SAML or LDAP configured, admin accounts in use, roles with search restrictions that are wider than intended.
  • TLS everywhere. Forwarder to indexer, search head to indexer, web interface. Self-signed certificates are common; expired ones are more common than they should be.
  • Monitoring Console. Whether it is configured, whether anyone looks at it, and what it has been saying.
  • Backups and change control. What would happen if the search head cluster captain disappeared tonight, and whether configuration lives in version control.
  • Upgrade path. Given all of the above, what the next safe upgrade looks like and what blocks it.

What we do with the list

The output is a short written report: what we found, what it costs the team today, and what we would fix first. It is deliberately ordered by impact rather than by severity label, because a platform with thirty medium findings and one small change that halves licence usage should start with that change.

If you want us to run this on your environment, an advisory engagement takes two to six weeks depending on size. If you would rather run it yourself, the list above is a fair start. Most of the searches behind it are in the Monitoring Console already.