Splunk Error Reference: Causes and Resolutions

This reference lists the Splunk Enterprise errors that administrators meet most often, with the usual cause and the steps to fix each one. It covers search, indexing, parsing, forwarding, HTTP Event Collector, licensing, clustering, authentication, startup and app deployment.

How to use this guide

Search this page for a short, distinctive phrase from your error rather than the whole line, because hostnames, paths and counts differ per environment. Messages follow the wording of Splunk Enterprise 9.x and later; older releases say “master” and “slave” where current ones say “manager” and “peer”. SPL searches and CLI commands named in the tables are written out in full in the Diagnostic toolkit at the end.

Where errors surface

LocationWhat it holdsHow to reach it
splunkd.logCore daemon errors: inputs, outputs, indexing, clustering, authentication$SPLUNK_HOME/var/log/splunk/splunkd.log, or search index=_internal sourcetype=splunkd
search.logErrors for one search job, including messages from each search peerJob Inspector, or $SPLUNK_HOME/var/run/splunk/dispatch/<sid>/search.log
metrics.logQueue fill, throughput and connection counts every 30 secondsindex=_internal source=*metrics.log
scheduler.logScheduled search runs, skips and their reasonsindex=_internal sourcetype=scheduler
license_usage.logDaily indexed volume by pool, index, sourcetype and host (on the license manager)index=_internal source=*license_usage.log
mongod.logKV store startup and replication errors$SPLUNK_HOME/var/log/splunk/mongod.log
web_service.logSplunk Web errors, including HTTP 500 pages$SPLUNK_HOME/var/log/splunk/web_service.log
Crash logsStack trace written when splunkd dies on a fatal signal$SPLUNK_HOME/var/log/splunk/crash-*.log
Messages menuBulletin banners in Splunk Web, including messages relayed from search peersMessages, top bar of Splunk Web
Health reportRed, yellow or green status per feature, with root cause textHealth icon beside your user name in Splunk Web
Monitoring ConsoleDashboards for indexing, search, resource use and forwarders across the deploymentSettings, then Monitoring Console

First-response routine

  1. Copy the exact message, its component name and its timestamp. A splunkd.log line reads: date, time, level (ERROR, WARN, INFO), component, then the message.
  2. Identify which instance wrote it. An error shown on a search head often originates on an indexer or forwarder.
  3. Establish what changed just before: a restart, an app install, a config push, a certificate renewal or a volume spike.
  4. Run splunk btool check --debug to rule out configuration typos.
  5. Check the host: free disk space, memory, CPU and open file limits.
  6. If the cause is still unclear, raise that component’s log level to DEBUG under Settings, Server settings, Server logging. Set it back afterwards.
  7. If you need Splunk Support, run splunk diag on the affected instance and attach the output to the case.

Search errors

Most search errors come from a concurrency or disk quota, a limit in limits.conf, or a search peer that stopped answering. They appear as banners above the results and in search.log.

Error messageCauseResolution
The maximum number of concurrent historical searches on this instance has been reachedThe instance is at its search concurrency ceiling. The default ceiling is max_searches_per_cpu x CPU cores + base_max_searches (1 x cores + 6), and scheduled searches get half of it.(1) List running jobs under Activity, Jobs, or in the Monitoring Console under Search, Activity. (2) Spread scheduled searches across the hour and give them a schedule window. (3) Shorten or rewrite long-running searches. (4) Add CPU cores or search heads. Raise the limits only if the CPUs have spare capacity.
This search could not be dispatched because the role-based concurrency limit of historical searches for role=<role> has been reached or ...the user-based quota... (usage=3, quota=3)The user or role has hit srchJobsQuota or cumulativeSrchJobsQuota in authorize.conf.(1) Ask the user to finish or delete their running jobs. (2) Check the role’s quotas under Settings, Roles, Resources. (3) Raise srchJobsQuota for that role if the workload justifies it.
The search job has been queued since the disk usage quota (500MB) for user '<user>' was reachedThe user’s stored search artifacts exceed the role’s srchDiskQuota.(1) Delete old jobs under Activity, Jobs. (2) Lower dispatch.ttl on saved searches that keep large results. (3) Raise srchDiskQuota for the role if needed.
Dispatch Runner: The number of search artifacts in the dispatch directory is higher than recommended (count=5123, warning threshold=5000)Too many job directories under var/run/splunk/dispatch. Usual sources are frequent scheduled searches with long time-to-live and alerts that keep artifacts after firing.(1) Find the searches creating most artifacts (Toolkit: dispatch count). (2) Reduce dispatch.ttl and alert expiry on them. (3) Move old artifacts out with splunk clean-dispatch /tmp/old-dispatch -24h@h.
Search not executed: The minimum free disk space (5000MB) reached for /opt/splunk/var/run/splunk/dispatchFree space on the dispatch partition is below minFreeSpace in server.conf.(1) Free space on the partition. (2) Clean the dispatch directory as above. (3) Do not lower minFreeSpace as the fix; a full disk also stops indexing.
Search auto-finalized after time limit (<n> seconds) reachedThe search ran past the role’s srchMaxTime, or a subsearch ran past maxtime (60 seconds by default). Results are partial.(1) Narrow the time range and add index= and sourcetype= filters. (2) Replace the subsearch or join with stats or a lookup. (3) Raise the limit only for roles that need long searches.
Subsearch produced <n> results, truncating to maxout 10000A subsearch returned more than maxout rows ([subsearch] in limits.conf). The outer search silently works on a truncated set.(1) Rewrite with stats, lookup or tstats so no subsearch is needed. (2) Filter the subsearch harder and return one field with fields and dedup. (3) Raise maxout only as a last resort; it cannot be set to 10,500 or above.
'stats' command: limit for values of field '<field>' reached. Some values may have been truncated or ignored.list() keeps 100 values per group by default, and values() can hit maxvalues when it is set.(1) Use values() or dc() instead of list() where order does not matter. (2) Split by an extra field to shrink each group. (3) Raise list_maxsize in [stats] if you need the full list.
Unknown search command '<name>'.The command is misspelled, the app that provides it is not installed on the search head, or the app is not shared globally.(1) Check the spelling. (2) Install the app on the search head (and on indexers for streaming commands). (3) Set the app’s permissions to Global or run the search inside that app.
Error in 'lookup' command: Could not construct lookup '<name>' or The lookup table '<name>' does not exist or is not available.The lookup file or definition is missing, private to another user or app, or excluded from the knowledge bundle sent to indexers.(1) Confirm the file and definition under Settings, Lookups. (2) Share both with the app or globally and grant read to the role. (3) If only indexers complain, add local=true to the lookup command or remove the file from the replication denylist.
Unable to distribute to peer named <peer> at uri https://<peer>:8089 because replication was unsuccessful.The knowledge bundle could not be copied to the indexer. The bundle is too large (often from big lookup files), or the transfer timed out.(1) Check bundle size in $SPLUNK_HOME/var/run/*.bundle on the search head. (2) Exclude large lookups and binaries with [replicationDenylist] in distsearch.conf. (3) If still too large, raise maxBundleSize on the search head and max_content_length on the peers. (4) Raise sendRcvTimeout on slow links.
Unknown error for peer <peer>. Search Results might be incomplete! or Timed out waiting for peer <peer>.The indexer is down, overloaded, or did not answer within the distributed search timeouts.(1) Check the peer’s status under Settings, Distributed search, Search peers. (2) Read the peer’s splunkd.log and crash logs for that time. (3) Test port 8089 from the search head. (4) If the peer is healthy but slow, raise connectionTimeout, sendTimeout and receiveTimeout in distsearch.conf.
Reached end-of-stream while waiting for more data from peer <peer>. Search results might be incomplete.The peer closed the connection mid-search: it restarted, its search process crashed, or a firewall or load balancer dropped the session.(1) Check whether the peer restarted or wrote a crash log. (2) Look for out-of-memory kills in the peer’s system log. (3) Check idle timeouts on network devices between search head and peer.
Search process did not exit cleanly, exit_code=<n>, description="exited with code <n>". Please look in search.log for this peer in the Job Inspector for more info.The search process on that instance crashed or was killed, most often by the Linux out-of-memory killer or Splunk’s own memory tracker.(1) Open search.log for the peer in the Job Inspector. (2) Run dmesg -T and look for oom-killer. (3) Reduce memory use: avoid large transaction, join and unbounded stats values(). (4) Add memory or set per-search limits (see Startup and resources).
status=skipped, reason="The maximum number of concurrent running jobs for this historical scheduled search on this instance has been reached" (in scheduler.log)The previous run of the same scheduled search is still running when the next one is due.(1) Compare the search’s run time with its interval (Toolkit: skipped searches). (2) Make it faster with a tighter time range, tstats or summary data. (3) Lengthen the interval so runs do not overlap.

Indexing and data ingestion errors

Indexing stops or slows when a queue in the pipeline fills, and the full queue furthest downstream is the bottleneck. Data moves through the parsing, aggregation, typing and index queues in that order, so check the index queue first and work backwards.

Error messageCauseResolution
group=queue, name=indexqueue, blocked=true, max_size_kb=500, current_size_kb=499 (in metrics.log)A pipeline queue is full. A full index queue points to slow disk writes or a blocked output. A full typing queue points to heavy regex in TRANSFORMS or SEDCMD. A full aggregation queue points to line merging and timestamp work.(1) Chart queue fill by queue name to find the furthest full queue (Toolkit: queue fill). (2) Index queue: check disk I/O wait and free space, and any forwarding output. (3) Typing queue: simplify or anchor the regexes named in props.conf and transforms.conf. (4) Aggregation queue: set SHOULD_LINEMERGE = false, LINE_BREAKER and TIME_FORMAT explicitly.
Now skipping indexing of internal audit events, because the downstream queue is not accepting data. Will keep dropping events until indexer congestion is remedied.The index queue is blocked, so Splunk drops its own audit events to cope. This is a symptom of the row above.(1) Check free disk space on index volumes. (2) Find the blocked queue as above. (3) If the instance forwards data, confirm the receiving indexers are reachable and not blocked.
The index processor has paused data flow. Too many tsidx files in idx=<index> bucket="<path>", waiting for the splunk-optimize indexing helper to catch up merging them. or Throttling indexer, too many tsidx files in bucket=<path>. Is splunk-optimize working?splunk-optimize cannot merge index files as fast as they are written. Usual reasons are slow storage, many hot buckets open at once, or bad timestamps spreading events across many buckets.(1) Measure disk throughput on the hot path; indexers need fast, low-latency storage. (2) Fix timestamp extraction for the noisy sourcetype. (3) Confirm splunk-optimize processes are running and not failing in splunkd.log. (4) If storage is healthy, raise maxRunningProcessGroups in indexes.conf.
The index processor has paused data flow. Current free disk space on partition '<mount>' has fallen to <n>MB, below the minimum of 5000MB. Data writes to index path '<path>' cannot safely proceed.Free space on an index partition is below minFreeSpace ([diskUsage] in server.conf). Indexing stops until space returns.(1) Free space or extend the volume. (2) Cap index size with maxTotalDataSizeMB and age with frozenTimePeriodInSecs. (3) Use volume stanzas with maxVolumeDataSizeMB so indexes cannot outgrow the disk. (4) Move non-index data, such as dispatch and core files, off the partition.
Received event for unconfigured/disabled/deleted index=<index> with source="..." host="..." sourcetype="...". So far received events from 1 missing index(es).An input sends to an index that does not exist or is disabled on the indexers. The events are dropped.(1) Create the index on every indexer, through the cluster bundle in a cluster. (2) Or correct index = in the forwarder’s inputs.conf. (3) Set lastChanceIndex in indexes.conf to catch misrouted events in future.
Failed to read size=<n> event(s) from rawdata in bucket='<bucket>' path='<path>'. Rawdata may be corrupt, see search.log. Results may be incomplete!The bucket is corrupt, usually after an unclean shutdown, a full disk or a storage fault.(1) Check the system log for disk errors first. (2) Stand-alone indexer: stop Splunk and run splunk fsck repair --one-bucket --bucket-path=<path>, or splunk rebuild <path>. (3) Cluster: remove the bad copy and let the manager replicate a good one.
Detected directory manually copied into its database, causing id conflicts [<path1>, <path2>].Two buckets in the same index share a bucket ID, typically after buckets were restored or copied by hand.(1) Stop Splunk. (2) Rename one bucket directory so its trailing ID is unused in that index. (3) Start Splunk and confirm the message clears.
ERROR BucketMover - coldToFrozenScript ... exited with non-zero statusThe archiving script fails, so buckets are never frozen and the disk keeps filling.(1) Run the script by hand as the Splunk user against a test bucket. (2) Fix the destination path, permissions or free space. (3) Consider coldToFrozenDir instead of a script.
The percentage of small buckets (<n>%) created over the last hour is high and exceeded the red thresholds (<n>%) for index=<index> (health report)Hot buckets roll too often. Causes are events with widely scattered timestamps, frequent restarts, or too few hot buckets for the data mix.(1) Find the sourcetype with bad or mixed timestamps and fix its props.conf. (2) Check for repeated restarts or forced rolls. (3) For indexes that legitimately mix old and new data, raise maxHotBuckets.
No error, but expected events are missing from searchThe events were indexed with the wrong timestamp, went to another index, or the role cannot search that index. Less often the input is disabled.(1) Search All time with index=* host=<host> to find where and when they landed, and compare _time with _indextime. (2) Check the role’s allowed indexes under Settings, Roles. (3) On the forwarder, run splunk list inputstatus and splunk btool inputs list --debug.

Parsing errors: timestamps and line breaking

Parsing warnings mean events are being stamped with the wrong time or split in the wrong place, and the fix is nearly always an explicit props.conf stanza for the sourcetype. Parsing settings take effect on the first full Splunk instance the data reaches, which is a heavy forwarder if you use one and the indexer otherwise. They apply only to data indexed after the change.

Error messageCauseResolution
WARN DateParserVerbose - Failed to parse timestamp in first MAX_TIMESTAMP_LOOKAHEAD (128) characters of event. Defaulting to timestamp of previous eventNo recognisable timestamp sits in the first 128 characters, or its format differs from what Splunk guessed.(1) Set TIME_PREFIX to a regex that ends just before the timestamp. (2) Set TIME_FORMAT to the exact strptime pattern. (3) Set MAX_TIMESTAMP_LOOKAHEAD to the timestamp’s length. (4) Test with the Add Data preview before deploying.
WARN DateParserVerbose - A possible timestamp match (<date>) is outside of the acceptable time window. If this timestamp is correct, consider adjusting MAX_DAYS_AGO and MAX_DAYS_HENCE.The parsed date is more than 2,000 days old or more than 2 days ahead. Often a number such as an ID or a port was misread as a date.(1) Anchor extraction with TIME_PREFIX and TIME_FORMAT so only the real timestamp matches. (2) If you are loading old data on purpose, raise MAX_DAYS_AGO for that sourcetype.
WARN DateParserVerbose - Time parsed (<date>) is too far away from the previous event's time (<date>) to be accepted. If this is a correct time, MAX_DIFF_SECS_AGO (3600) or MAX_DIFF_SECS_HENCE (604800) may be overly restrictive.Consecutive events in one source jump more than an hour back or a week forward. Day and month may be swapped, or the file mixes events out of order.(1) Set TIME_FORMAT explicitly to remove day and month ambiguity. (2) If the source really is out of order, raise MAX_DIFF_SECS_AGO and MAX_DIFF_SECS_HENCE.
WARN DateParserVerbose - Accepted time (<date>) is suspiciously far away from the previous event's time (<date>), but still accepted because it was extracted by the same pattern.Same pattern as above, but Splunk kept the time. Usually harmless for sparse logs and a warning sign for busy ones.(1) Spot-check events from that source for wrong _time. (2) If wrong, apply the explicit timestamp settings above.
No error, but events are offset by a whole number of hoursThe timestamp carries no time zone, so Splunk applies the zone of the forwarder or indexer.(1) Set TZ in props.conf for the host, source or sourcetype. (2) Prefer logging in UTC or with an explicit offset at the source.
WARN LineBreakingProcessor - Truncating line because limit of 10000 bytes has been exceeded with a line length >= <n>A single line is longer than TRUNCATE (10,000 bytes by default). Either the events really are long, or LINE_BREAKER never matches and several events run together.(1) Check a raw sample to see which case applies. (2) Fix LINE_BREAKER first if events are running together. (3) Raise TRUNCATE for that sourcetype if events are legitimately long. Avoid TRUNCATE = 0 (unlimited).
WARN AggregatorMiningProcessor - Breaking event because limit of 256 has been exceededA multi-line event reached MAX_EVENTS (256 lines) while line merging was on.(1) Set SHOULD_LINEMERGE = false and a LINE_BREAKER that matches the start of each event. This is faster and more reliable than merging. (2) If you must merge, raise MAX_EVENTS.
WARN AggregatorMiningProcessor - Changing breaking behavior for event stream because MAX_EVENTS (256) was exceeded without a single event break. Will set BREAK_ONLY_BEFORE_DATE to FalseLine merging found no event boundary at all, typically because the events have no timestamp where Splunk expects one.(1) Define LINE_BREAKER and SHOULD_LINEMERGE = false for the sourcetype. (2) Add TIME_PREFIX and TIME_FORMAT, or DATETIME_CONFIG = CURRENT if the data has no timestamp.
ERROR JsonLineBreaker - JSON StreamId:<n> had parsing error: Unexpected character while looking for valueINDEXED_EXTRACTIONS = json is set but the content is not valid JSON, for example a text prefix before the object or a truncated record.(1) Validate a raw sample. (2) Remove non-JSON prefixes at the source, or drop INDEXED_EXTRACTIONS and use search-time KV_MODE = json. (3) Remember that structured parsing runs on the universal forwarder, so put this props.conf there.
Events arrive with a sourcetype ending in -too_small or a number such as access-2No sourcetype was assigned, so Splunk guessed one from the file content.(1) Set sourcetype = in the inputs.conf stanza. (2) Add a matching props.conf stanza with the six core settings: LINE_BREAKER, SHOULD_LINEMERGE, TIME_PREFIX, TIME_FORMAT, MAX_TIMESTAMP_LOOKAHEAD, TRUNCATE.

Forwarder and network errors

Forwarder errors fall into two groups: the forwarder cannot reach or hand data to an indexer, or it cannot read the source in the first place. Read splunkd.log on the forwarder itself; little of this reaches the indexer.

Error messageCauseResolution
WARN TcpOutputProc - Cooked connection to ip=<ip>:9997 timed outThe forwarder cannot open a connection to the indexer. A firewall blocks the port, the indexer is down, or the address or port in outputs.conf is wrong.(1) On the forwarder, run splunk list forward-server to see active and inactive targets. (2) Test the port with nc -vz <indexer> 9997. (3) On the indexer, confirm receiving is enabled with splunk display listen or a [splunktcp://9997] stanza. (4) Open the port on host and network firewalls.
ERROR TcpOutputFd - Connection to host=<ip>:9997 failed or Connection refusedThe host is reachable but nothing listens on that port.(1) Enable receiving on the indexer under Settings, Forwarding and receiving. (2) Check the indexer is running and finished starting. (3) Confirm the forwarder uses the receiving port, not the management port 8089.
WARN TcpOutputProc - Forwarding to indexer group <group> blocked for <n> seconds. or The TCP output processor has paused the data flow. Forwarding to host_dest=<host> inside output group <group> has been blocked for blocked_seconds=<n>.Every indexer in the group is unreachable or refusing data because its own queues are full. The forwarder holds data and may stop reading files.(1) Check indexer queues and disk first (see Indexing and data ingestion errors). (2) Check connectivity as in the rows above. (3) List all indexers in outputs.conf so load balancing can route around a slow one. (4) Add indexer capacity if the block follows volume peaks.
WARN TailReader - Could not send data to output queue (parsingQueue), retrying...The forwarder’s own queues are full because its output is blocked or throttled.(1) Resolve the output block above. (2) Check for the throughput limit in the next row.
INFO ThruputProcessor - Current data throughput (256 kb/s) has reached maxKBps. As a result, data forwarding may be throttled.The universal forwarder’s default rate limit of 256 KBps is too low for the host’s log volume. Data arrives late.(1) Set maxKBps under [thruput] in the forwarder’s limits.conf to a higher value, or 0 for no limit. (2) Restart the forwarder. (3) Watch indexer load after lifting limits across many hosts.
ERROR TcpInputProc - Message rejected. Received unexpected message of size=<n> bytes from src=<ip>:<port> in streaming mode. Maximum message size allowed=67108864.Something other than a Splunk forwarder is sending to the splunktcp port: raw syslog, an HTTP client, or a forwarder using TLS against a plain-text port.(1) Identify the sender from src. (2) Send raw or syslog data to a [tcp://<port>] or [udp://<port>] input instead. (3) Make TLS settings match on both ends.
ERROR TcpInputProc - Error encountered for connection from src=<ip>:<port>. error:1408F10B:SSL routines:ssl3_get_record:wrong version numberA forwarder is sending plain text to a TLS-enabled port.(1) Configure TLS in the forwarder’s outputs.conf. (2) Or point it at the non-TLS receiving port if one is intended.
ERROR TailReader - File will not be read, seekptr checksum did not match (file=<path>). Last time we saw this initcrc, filename was different. You may wish to use larger initCrcLen for this sourcetype, or a CRC salt on this source.The first 256 bytes match a file already read, so Splunk treats the new file as a rotated copy. Common with files that start with an identical header.(1) Raise initCrcLength in the inputs.conf stanza so the checksum covers past the header. (2) Use crcSalt = <SOURCE> only for files that are never renamed or rotated, otherwise it re-indexes them.
WARN TailReader - Insufficient permissions to read file='<path>' (hint: Permission denied, UID: <n>, GID: <n>)The account running the forwarder cannot read the file.(1) Add the forwarder’s user to the group that owns the logs, or grant read with setfacl. (2) Check every parent directory is traversable. (3) Make sure log rotation recreates files with the same permissions.
INFO TailReader - File descriptor cache is full (100), trimming...The forwarder monitors more files than it can hold open. Frequent messages mean slow pickup of new data.(1) Narrow the monitor path and add an allow or deny list. (2) Set ignoreOlderThan to skip stale files. (3) Archive or delete old logs. (4) Raise max_fd under [inputproc] in limits.conf.
ERROR ExecProcessor - message from "<script path>" ... or Couldn't start command "<script path>": Permission deniedA scripted input failed: the file is not executable, its interpreter is missing, or the script wrote an error.(1) Make the script executable. (2) Run it as the Splunk user with splunk cmd <script path> and read the output. (3) Check the app supports the Python version on that instance.
Bind error on UDP or TCP port 514: Permission deniedPorts below 1024 need root, and Splunk should not run as root.(1) Listen on a high port such as 5514 and redirect 514 to it with a firewall rule. (2) Better, receive syslog with a dedicated syslog server and have a forwarder monitor its files.
WARN AutoLoadBalancedConnectionStrategy - Possible duplication of events with channel=<id>, streamId=<n>, offset=<n> on host=<ip>:9997With useACK = true, the forwarder did not get an acknowledgement in time and sent the block again.(1) Check the named indexer for restarts or blocked queues at that time. (2) Occasional messages are the price of acknowledgement; frequent ones mean an unhealthy indexer or network path.
No error, but events are duplicatedFiles were re-read after rotation (crcSalt = <SOURCE> on rotating logs, or copy-and-truncate rotation), or two inputs cover the same path.(1) Remove crcSalt from rotating sources. (2) Run splunk btool inputs list --debug to find overlapping monitor stanzas. (3) Check the same host is not sending through two forwarders.

HTTP Event Collector errors

HEC answers every request with a JSON body such as {"text":"Invalid token","code":4}, and the code tells you which of the rows below applies. Codes and messages are from Splunk’s Troubleshoot HTTP Event Collector page; codes 0 (Success) and 17 (HEC is healthy) are normal responses and are left out.

CodeHTTP statusMessageCauseResolution
1403Token disabledThe token exists but is disabled, or HEC is disabled globally.(1) Enable the token under Settings, Data inputs, HTTP Event Collector. (2) Check Global Settings has All Tokens set to Enabled.
2401Token is requiredThe request has no Authorization header.Send Authorization: Splunk <token> with every request.
3401Invalid authorizationThe header is malformed, for example Bearer <token> or a missing Splunk prefix.(1) Use the exact form Authorization: Splunk <token>. (2) With basic authentication, put the token in the password field.
4403Invalid tokenThe token value does not exist on the instance that answered. Often one node behind a load balancer lacks the token.(1) Compare the token value with the one shown in Splunk Web. (2) Deploy the same HEC inputs.conf to every receiving instance.
5400No dataThe request body is empty.Check the client builds and attaches the payload before sending.
6400Invalid data formatThe body is not valid JSON for the event endpoint: single quotes, a trailing comma, raw text, or misspelled metadata keys.(1) Validate the JSON. (2) Wrap the payload as {"event": ...}. (3) Send plain text to /services/collector/raw instead. (4) Read HttpInputDataHandler lines in splunkd.log for the parse position.
7400Incorrect indexThe event names an index the token is not allowed to write to, or one that does not exist.(1) Add the index to the token’s allowed indexes. (2) Create the index on the indexers. (3) Or drop the index key to use the token’s default.
8500Internal server errorsplunkd failed while handling the request.(1) Read splunkd.log on the receiver for that timestamp. (2) Retry. (3) Open a Support case if it persists.
9503Server is busyThe receiver cannot take more data, usually because indexing queues are full.(1) Resolve blocked queues (see Indexing and data ingestion errors). (2) Make the client retry with backoff. (3) Add HEC receivers behind the load balancer.
10400Data channel is missingIndexer acknowledgement is on for the token, or the raw endpoint is used, and no channel was sent.Send an X-Splunk-Request-Channel: <GUID> header or a ?channel=<GUID> parameter.
11400Invalid data channelThe channel value is not a valid GUID.Generate one GUID per client and reuse it.
12400Event field is requiredThe JSON object has no event key.Add the event key; metadata alone is not an event.
13400Event field cannot be blankThe event value is empty or null.Filter out empty records in the client.
14400ACK is disabledThe client polls /services/collector/ack but the token does not use acknowledgement.(1) Enable indexer acknowledgement on the token. (2) Or stop the client polling for acknowledgements.
15400Error in handling indexed fieldsThe fields key is malformed. It must be a flat object, not nested objects.Flatten fields to simple key and value pairs, or move nested data into event.
16400Query string authorization is not enabledThe token was passed as ?token= but query string authentication is off.(1) Send the token in the header. (2) Or set allowQueryStringAuth = true in the token’s stanza.
18503HEC is unhealthy, queues are fullThe health endpoint reports blocked indexing queues.Treat as code 9. Have the load balancer take the node out until it recovers.
19503HEC is unhealthy, ack service unavailableAcknowledgement tracking has reached its limits, usually because clients never collect their acknowledgements.(1) Make clients poll the ack endpoint and reuse channels. (2) Review the ack limits under [http_input] in limits.conf.
20503HEC is unhealthy, queues are full, ack service unavailableBoth conditions above at once.Apply the fixes for codes 18 and 19.
21400Invalid tokenSame condition as code 4, returned with HTTP 400.As code 4.
22400Token disabledSame condition as code 1, returned with HTTP 400.As code 1.
23503Server is shutting downThe receiver is stopping or restarting.(1) Retry with backoff. (2) Drain the node from the load balancer before planned restarts.
24200HEC queue is approaching its capacity limitThe request was accepted, but the queue is nearly full.Slow the sender or add receiving capacity before requests start failing.
25200HEC ACK is approaching its capacity limitThe request was accepted, but acknowledgement tracking is nearly full.Poll acknowledgements more often and reuse channels.
26429HEC queue is at capacity and cannot process any more requestsThe queue is full and the request was rejected.(1) Back off and resend. (2) Fix blocked indexing queues. (3) Add receivers.
27429HEC ACK channel is at capacity and cannot process any more requestsThe acknowledgement channel is full and the request was rejected.(1) Collect outstanding acknowledgements. (2) Back off and resend.

Some failures happen before HEC returns a code:

SymptomCauseResolution
Connection refused on port 8088HEC is not enabled on that instance, or it listens on another port.(1) Enable HEC in Global Settings. (2) Confirm the port there and in the firewall.
curl: (35) ... wrong version numberThe client uses https while HEC has SSL turned off, or the reverse.Match the URL scheme to the Enable SSL setting.
curl: (60) SSL certificate problem: self-signed certificateHEC presents Splunk’s default certificate.(1) Install a CA-signed certificate for HEC. (2) Use -k only for testing.
HTTP 404 Not FoundThe URL path is wrong.Use /services/collector/event for JSON events or /services/collector/raw for raw text.

Licensing errors

A license warning is recorded for each calendar day your indexed volume exceeds the licensed amount, and enough warnings block search while indexing continues. The thresholds below are from Splunk’s About license violations page for version 10.4.

License typeWhen search is blocked
Enterprise, 100 GB per day or moreNever. Warnings are issued but search stays on.
Enterprise, under 100 GB per day45 warnings in a rolling 60-day window. A reset license from Splunk Sales re-enables search.
Enterprise infrastructure (vCPU)Does not violate.
Trial, Dev/Test, Developer5 or more warnings in a rolling 30-day window.
Free3 or more warnings in a rolling 30-day window. No reset license is available.

Searches against internal indexes such as _internal keep working during a violation, so you can still diagnose the problem.

Error messageCauseResolution
This pool has exceeded its configured poolsize=<n> bytes. A warning has been recorded for all members.Today’s indexed volume passed the pool’s quota. The warning counts once midnight passes on the license manager.(1) Open Settings, Licensing, Usage report and split by sourcetype, host and index to find the spike (Toolkit: license usage). (2) Decide whether it is a one-off, such as debug logging left on, or a new baseline. (3) Move spare quota from another pool, or add license volume. (4) For unwanted data, filter at the forwarder or route it to nullQueue.
Error in 'litsearch' command: Your Splunk license expired or you have exceeded your license limit too many times. Renew your Splunk license by visiting www.splunk.com/store or calling 866.GET.SPLUNK.The instance is in violation or its license has expired, so search is blocked.(1) Check the warning count and expiry under Settings, Licensing. (2) Install the renewed license if it expired. (3) If in violation on an Enterprise license, request a reset license and install it on the license manager. (4) Fix the over-ingestion so the warnings do not return.
LMTracker errors containing failed to send rows or unable to connectA license peer cannot reach the license manager. After 72 hours without contact the peer goes into violation and search is blocked on it.(1) Test port 8089 from the peer to the manager. (2) Check manager_uri under [license] in the peer’s server.conf. (3) Confirm pass4SymmKey matches on both sides. (4) Search index=_internal LMTracker error to confirm it has cleared.
License peers appear and disappear on the manager, or usage is under-reportedTwo or more instances share one GUID because they were cloned from an image that had already been started.(1) On each clone, stop Splunk and run splunk clone-prep-clear-config, then start it. (2) Build future images from an instance that has never been started.
Trial license expired and login shows only the Free optionThe 60-day trial ended.(1) Install an Enterprise license. (2) Or accept the Free license, noting it removes authentication, alerting and distributed search.

Indexer clustering errors

Most indexer cluster alarms mean the manager cannot keep the required number of bucket copies, either because a peer is unavailable or because copies cannot be made. Start on the cluster manager under Settings, Indexer clustering, and with splunk show cluster-status --verbose.

Error messageCauseResolution
Search Factor is Not Met, Replication Factor is Not Met, All Data is not SearchableSome buckets have fewer copies than the cluster’s replication factor or search factor requires. A peer is down or restarting, fix-up tasks are still running, or there are too few peers.(1) On the manager, open Indexer clustering, Indexes, Bucket Status to see pending fix-up tasks and their reasons. (2) Bring missing peers back. (3) Wait for fix-ups to finish after a restart; this is normal for a while. (4) Confirm the number of healthy peers is at least the replication factor, per site in a multisite cluster. (5) Check peers have free disk space.
Cannot fix search count as the bucket hasn't rolled yet.A hot bucket lacks a searchable copy, and the manager cannot fix it until the bucket rolls to warm.(1) Wait; it normally clears when the bucket rolls. (2) To clear it now, roll the bucket from Bucket Status using the Action menu, or roll the index’s hot buckets on the originating peer.
Missing enough suitable candidates to create searchable copy in order to meet replication policy. Missing={ default:1 }No eligible peer can take another copy. Peers are down, in detention, full, or the site lacks enough peers for the site replication settings.(1) Check each peer’s status on the manager’s Peers tab. (2) Free disk space on peers in automatic detention. (3) Add peers, or lower the replication or search factor to match the peers you have.
Too many bucket replication errors to target peer=<ip>:<port>. Will stop streaming data to this target while errors persist.The source peer cannot stream buckets to the target on the replication port. The port is blocked, the target is overloaded or out of disk, or TLS settings differ.(1) Test the replication port from source to target. (2) Read the target’s splunkd.log for the matching time. (3) Check disk space and I/O on the target. (4) Put the target in manual detention while you repair it.
Peer status AutomaticDetentionFree disk on the peer fell below minFreeSpace. The peer stops indexing and stops receiving replicated data.(1) Free or add disk space on the peer. (2) Review index size and retention settings so it does not recur.
Failed to register with cluster master reason: failed method=POST path=/services/cluster/master/peers ...The peer cannot join the cluster. pass4SymmKey differs from the manager’s, the manager is unreachable, or the peer runs an incompatible version.(1) Set the same plain-text pass4SymmKey under [clustering] on the peer and restart it. (2) Test port 8089 to the manager. (3) Check manager_uri on the peer. (4) Confirm the manager’s version is the same as or newer than the peer’s.
splunk apply cluster-bundle fails with a bundle validation errorA file under manager-apps (older: master-apps) is invalid. Typical faults are an index path that does not exist on peers, an undefined volume, or a typo in a stanza.(1) Run splunk validate cluster-bundle --check-restart. (2) Run splunk show cluster-bundle-status to read the per-peer error. (3) Correct the file and apply again.
The searchhead is unable to update the peer information. Error = 'Unable to reach the cluster manager' for manager=https://<host>:8089The search head cannot contact the cluster manager, so its list of peers may be stale.(1) Check the manager is running. (2) Test port 8089 from the search head. (3) Check manager_uri and pass4SymmKey in the search head’s [clustering] stanza.
Rolling restart or bundle push appears stuckThe manager waits for each peer to return before restarting the next, and one peer has not come back.(1) Find the peer that is down in splunk show cluster-status. (2) Read its splunkd.log for the startup failure, often a config error from the new bundle. (3) Fix and start it; the restart then continues.

Search head clustering and KV store errors

Search head cluster faults trace back to three things: members that cannot agree on a captain, a member whose configuration has drifted from the captain’s, or a KV store that will not start or sync. Run splunk show shcluster-status and splunk show kvstore-status on any member first.

Error messageCauseResolution
KV Store initialization failed. Please contact your system administrator., Failed to start KV Store process. See mongod.log and splunkd.log for details. or KV Store changed status to failed.mongod cannot start. Common reasons are an expired server certificate, wrong file ownership under var/lib/splunk/kvstore, a full disk, or a stale lock after an unclean shutdown.(1) Read mongod.log for the first error after startup. (2) Check certificate expiry with openssl x509 -enddate -noout -in $SPLUNK_HOME/etc/auth/server.pem. (3) Reset ownership of $SPLUNK_HOME to the Splunk user. (4) Free disk space. (5) Last resort on a cluster member: splunk clean kvstore --local, then restart so it resyncs. On a stand-alone instance this erases KV data, so run splunk backup kvstore first.
Error in 'inputlookup' command: External command based lookup '<name>' is not available because KV Store initialization has failed.A search uses a KV store lookup while the KV store is down.Fix the KV store as in the row above.
KV store status shows a member as stale or out of syncThe member was offline long enough to fall behind the replication log.(1) Run splunk show kvstore-status to identify the stale member. (2) Run splunk resync kvstore on it. (3) If that fails, stop it, run splunk clean kvstore --local, and start it so it copies from the others.
Members report no captain, or splunk show shcluster-status fails with a captain errorThe cluster lost its majority, so no captain can be elected. Members are down, cannot reach each other, or disagree on pass4SymmKey.(1) Bring members back until more than half are up. (2) Test port 8089 and the replication port between members. (3) Check clocks are in sync. (4) If a majority is gone for good, set a static captain as a temporary measure and revert once members are restored.
Error pulling configurations from the search head cluster captain; consider performing a destructive configuration resync on this search head cluster member.The member’s replicated configuration no longer shares a baseline with the captain, usually after long downtime.(1) Run splunk resync shcluster-replicated-config on that member. (2) Expect it to discard local changes made on that member since it diverged.
ConfReplicationException: Error pushing configurations to captain=<uri>, consecutiveErrors=<n>Same drift as above, seen from the member trying to send its changes.Resync the member as in the row above.
splunk apply shcluster-bundle fails with an authentication or connection errorThe deployer cannot reach the member given in -target, or the deployer’s pass4SymmKey under [shclustering] differs from the members’.(1) Point -target at a running member’s management URI. (2) Set the same plain-text [shclustering] key on the deployer and restart it. (3) Check the admin credentials passed with -auth.
Scheduled searches run twice or not at all in the clusterCaptaincy keeps changing, so scheduling is interrupted or repeated around each change.(1) Check captain changes in splunkd.log (search for SHCRaftConsensus). (2) Stabilise the network between members. (3) Confirm all members run the same Splunk version.

Authentication, TLS and permission errors

Login and trust failures come from four sources: the identity provider, a certificate, a shared secret that differs between instances, or a role that lacks a capability. Failed logins are recorded in the _audit index, and TLS errors in splunkd.log under SSLCommon and X509Verify.

Error messageCauseResolution
Login failed in Splunk WebWrong password, a locked account after repeated failures, or an external user with no Splunk role.(1) Search index=_audit action="login attempt" info=failed for the user and reason. (2) Unlock or reset the user under Settings, Users. (3) If the only admin password is lost, stop Splunk, move etc/passwd aside, create etc/system/local/user-seed.conf with a new admin user, and start Splunk.
Remote login has been disabled for 'admin' with the default password. Either set the password, or override by changing the 'allowRemoteLogin' setting in your server.conf file.The admin account on the instance, usually a universal forwarder, has no password set or still has the old default one, so remote management calls are refused.(1) Seed an admin password with etc/system/local/user-seed.conf and restart, or set one with splunk edit user admin -password <new>. (2) Do not set allowRemoteLogin = always.
ERROR ScopedLDAPConnection - strategy="<name>" Error binding to LDAP. reason="Can't contact LDAP server"Splunk cannot reach the directory server, or the LDAPS handshake fails.(1) Test the host and port (389 or 636) from the search head. (2) For LDAPS, add the directory’s CA certificate to $SPLUNK_HOME/etc/openldap/ldap.conf. (3) Check the host name in the strategy matches the certificate.
Error binding to LDAP. reason="Invalid credentials"The bind account’s DN or password is wrong, expired or locked.(1) Test the bind DN and password with ldapsearch. (2) Update the password in the LDAP strategy. (3) Use a service account whose password does not expire unannounced.
ERROR AuthenticationProviderLDAP - Could not find user="<user>" with strategy="<name>"The user is outside userBaseDN, the user name attribute is wrong, or the user belongs to no group mapped to a Splunk role.(1) Confirm the user’s DN falls under userBaseDN. (2) Check userNameAttribute (often sAMAccountName for Active Directory). (3) Map at least one of the user’s groups to a role.
SAML login fails with a message that no valid Splunk role was found, or that the response contains no group informationThe identity provider sends no role or group attribute, or none of the groups sent is mapped to a Splunk role.(1) Capture the assertion with a SAML tracer and check the group attribute. (2) Add the attribute in the identity provider. (3) Map the group names under Settings, Authentication methods, SAML Groups.
SAML login fails with a signature verification errorThe identity provider’s signing certificate changed and Splunk still holds the old one.(1) Re-import the identity provider’s metadata. (2) Or replace the certificate referenced by idpCertPath. (3) Reload authentication or restart.
ERROR SSLCommon - Can't read certificate file <path> errno=<n> error:0906D06C:PEM routines:PEM_read_bio:no start lineThe certificate file is missing, unreadable by the Splunk user, not in PEM format, or assembled in the wrong order.(1) Check the path and file ownership. (2) Build the file in this order: server certificate, private key, then intermediate and root certificates. (3) Check sslPassword matches the private key’s passphrase.
WARN X509Verify - X509 certificate (CN=<name>) failed validation; error=10, reason="certificate has expired"A certificate in the chain has expired.(1) Find the expiry with openssl x509 -enddate -noout -in <cert>. (2) Renew and replace it on every instance that uses it. (3) Restart. (4) Add certificate expiry to your monitoring.
Validation fails with reason="unable to get local issuer certificate" or reason="self signed certificate in certificate chain"The issuing CA is not in the trust store of the instance doing the verification.(1) Add the CA chain to the file named by sslRootCAPath in server.conf. (2) Restart. (3) Test with openssl s_client -connect <host>:<port> -CAfile <ca file>.
sslv3 alert handshake failure, no shared cipher or tlsv1 alert protocol versionThe two ends share no TLS version or cipher suite. Typical after hardening indexers while old forwarders remain.(1) Compare sslVersions and cipherSuite on both ends with btool. (2) Upgrade old forwarders. (3) As a stopgap, allow a common version on the receiving side.
HTTP 401 or call not properly authenticated between Splunk instancespass4SymmKey differs between the two instances for that feature (clustering, search head clustering, licensing or deployer).(1) Set the same plain-text key in the matching stanza on each instance. (2) Restart each one so the key is re-encrypted locally. (3) Never copy an encrypted $7$ value between instances unless they share splunk.secret.
Failed to decrypt errors after copying configuration between instancesAn encrypted password was copied from an instance with a different etc/auth/splunk.secret.(1) Replace the encrypted value with the plain-text password and restart. (2) Or deploy a common splunk.secret before first start on new instances.
You do not have permission to perform this operation or [HTTP 403] Client is not authorized to perform requested actionThe user’s roles lack the required capability, or the object is not shared with the user.(1) Identify the capability needed, for example schedule_search or edit_tcp. (2) Add it to an appropriate role under Settings, Roles. (3) Check the object’s sharing and read or write permissions.

Startup, configuration and resource errors

When Splunk will not start or keeps dying, the cause is nearly always a port conflict, a configuration typo, file ownership, or an operating system limit. Start-up messages print to the terminal and to splunkd.log; crashes leave a crash-*.log file.

Error messageCauseResolution
ERROR: The mgmt port [8089] is already bound. Splunk needs to use this port. or ERROR: The http port [8000] is already bound.Another process holds the port: a leftover splunkd, a second Splunk install, or another application.(1) Find the owner with ss -ltnp and the port number. (2) Stop the stale process. (3) If another application needs the port, change Splunk’s with splunk set splunkd-port <n> or splunk set web-port <n>.
Invalid key in stanza [<stanza>] in <file>, line <n>: <key> (value: <value>).A setting is misspelled, placed in the wrong file or stanza, or no longer exists in this version. Splunk ignores the line.(1) Run splunk btool check --debug for the full list. (2) Compare the key with the matching .spec file in $SPLUNK_HOME/etc/system/README. (3) Correct or remove the line.
Your indexes and inputs configurations are not internally consistent. For more information, run 'splunk btool check --debug'An input points to an index that is not defined on this instance, or indexes.conf has conflicting paths.(1) Run the suggested command. (2) Define the missing index or correct the input. On a forwarder that only forwards, the warning is harmless.
homePath='<path>' of index=<index> on unusable filesystem.The index path does not exist, is not mounted, is not writable by the Splunk user, or sits on a file system Splunk cannot lock files on.(1) Check the mount is present. (2) Create the directory and give it to the Splunk user. (3) Move indexes off unsupported network file systems.
Couldn't open "<path>": Permission denied at start-upSplunk was once started as root, leaving root-owned files that the service account cannot write.(1) Stop Splunk. (2) Run chown -R splunk:splunk $SPLUNK_HOME (substitute your service account). (3) Start Splunk as that account. (4) Enable boot-start with -user so it never starts as root again.
Too many open files, with start-up lines showing Limit: open files: 1024The operating system’s file descriptor limit for the Splunk process is too low.(1) Under systemd, set LimitNOFILE=64000 and LimitNPROC=16000 in the Splunk unit file and reload. (2) Under init scripts, set the same in /etc/security/limits.conf. (3) Restart and check the ulimit lines in splunkd.log.
Start-up or health check warning that transparent huge pages (THP) are enabledTransparent huge pages are enabled in Linux, which degrades Splunk’s memory performance.(1) Disable them at boot with the kernel option transparent_hugepage=never or a start-up unit. (2) Reboot or restart Splunk. (3) Verify in /sys/kernel/mm/transparent_hugepage/enabled.
Waiting for web server at http://127.0.0.1:8000 to be available... WARNING: web interface does not seem to be available!splunkd started but Splunk Web did not, or splunkd died soon after start.(1) Run splunk status. (2) Read the end of splunkd.log and web_service.log. (3) Look for a new crash log. (4) Check the web certificate settings in web.conf if TLS is enabled for Splunk Web.
splunkd disappears with nothing in splunkd.log; system log shows Out of memory: Killed process <pid> (splunkd)The Linux out-of-memory killer ended Splunk. Causes are memory-heavy searches, too many concurrent searches, or very large lookups.(1) Confirm with dmesg -T. (2) Find the heavy searches (Toolkit: memory by search). (3) Enable enable_memory_tracker with a per-search threshold under [search] in limits.conf. (4) Reduce search concurrency or add memory.
Crash log containing Received fatal signal 11 (Segmentation fault) or signal 6 (Aborted)A software defect, a corrupt bucket or file, or failing hardware.(1) Note the crashing thread and backtrace in the crash log. (2) Check the release notes’ known issues for your version. (3) Upgrade to the latest maintenance release. (4) Send a splunk diag and the crash log to Support.
500 Internal Server Error in Splunk WebA Splunk Web or app-side Python error, often after an app install or upgrade.(1) Read web_service.log and python.log for the traceback. (2) Disable the app named in the traceback and retest. (3) Clear the browser cache after upgrades.
File Integrity checks found <n> files that did not match the system-provided manifest.Files shipped with Splunk were edited, or an upgrade left old files behind.(1) Run splunk validate files to list them. (2) Move custom changes out of default directories into local. (3) Reinstall the same version over the top to restore shipped files.
SyntaxError: Missing parentheses in call to 'print' from an app scriptThe app’s scripts were written for Python 2, which current Splunk versions no longer include.(1) Update the app from Splunkbase. (2) Run the Upgrade Readiness App to list other affected apps. (3) Port private scripts to Python 3.

Deployment server and app errors

When a forwarder does not receive its apps, either it is not reaching the deployment server or no server class matches it. Check splunkd.log on the client for DeploymentClient lines and Settings, Forwarder management on the server.

Error messageCauseResolution
WARN DC:DeploymentClient - Unable to send handshake message to deployment server. Error status is: not_connected or DC:PhonehomeThread - Attempted handshake <n> times.The client cannot reach the deployment server. The targetUri is wrong, port 8089 is blocked, or the server is down or overloaded.(1) On the client, run splunk show deploy-poll and splunk btool deploymentclient list --debug. (2) Test port 8089 to the server. (3) Check the server is running. (4) For large fleets, raise phoneHomeIntervalInSecs to reduce load on the server.
Client connects but never appears, or several hosts appear as one clientHosts were cloned from an image that had already been started, so they share a GUID.(1) On each clone, stop Splunk, run splunk clone-prep-clear-config, and start it. (2) Rebuild the image from an instance that has never been started.
Client appears but receives no appsNo server class matches the client. The allow list uses a host name while the server sees an IP address or fully qualified name, or a machine type filter excludes it.(1) In Forwarder management, read the client’s host name, IP and client name exactly as the server sees them. (2) Adjust the allow list in serverclass.conf to match. (3) Run splunk reload deploy-server.
App changed on the server but clients keep the old versionThe server has not re-read its deployment-apps directory.(1) Run splunk reload deploy-server. (2) Confirm the app’s checksum changed in Forwarder management. (3) Check the app is set to restart the client if the change needs it.
WARN DeployedApplication - Installing app=<app> to='<path>' failed or a checksum mismatchThe download was cut short, the client’s disk is full, or the Splunk user cannot write to etc/apps.(1) Check free disk on the client. (2) Fix ownership of $SPLUNK_HOME/etc/apps. (3) Check for proxies or firewalls truncating large transfers.
Unable to initialize modular input "<name>" defined in the app "<app>": Introspecting scheme=<name>: script running failed (exited with code 1).The app’s input script cannot run on this instance: a missing library, an unsupported Python version, or wrong file permissions.(1) Run the script by hand with splunk cmd python <script> --scheme and read the error. (2) Check the app’s compatibility with your Splunk version. (3) Disable the input on instances that do not need it, such as indexers.
App install from Splunkbase fails with a connection or timeout errorThe instance has no route to the internet, or needs a proxy.(1) Configure [proxyConfig] in server.conf. (2) Or download the package elsewhere and use Install app from file.
A setting changed in an app has no effectAnother file with higher precedence sets the same key. etc/system/local beats every app, and local beats default.(1) Run splunk btool <conf name> list <stanza> --debug to see which file wins. (2) Remove or correct the overriding line. (3) Restart or reload as that setting requires.

Diagnostic toolkit

These are the commands and searches the tables refer to. Run CLI commands from $SPLUNK_HOME/bin as the Splunk service account, and run the searches on a search head or the Monitoring Console instance.

CLI commands

CommandWhat it tells you
splunk btool check --debugConfiguration typos, invalid keys and stanzas across all apps
splunk btool <conf name> list --debugThe effective value of every setting and the file it comes from
splunk statusWhether splunkd and its helpers are running
splunk list forward-serverOn a forwarder: configured indexers, and which are active
splunk list inputstatusOn a forwarder: each monitored file, how far it has been read, and why a file is skipped
splunk display listenOn an indexer: the port receiving forwarder data
splunk show cluster-status --verboseOn a cluster manager: peer state, replication and search factor status
splunk show cluster-bundle-statusOn a cluster manager: bundle validation and push status per peer
splunk show shcluster-statusOn a search head cluster member: captain, member state, last heartbeat
splunk show kvstore-statusKV store state and replication status per member
splunk show deploy-pollOn a deployment client: the deployment server it polls
splunk reload deploy-serverOn a deployment server: re-read server classes and apps
splunk validate filesShipped files that differ from the install manifest
splunk diagBuilds a diagnostic archive for Splunk Support

Errors and warnings by component

index=_internal sourcetype=splunkd (log_level=ERROR OR log_level=WARN) earliest=-24h
| stats count by host, component, log_level
| sort - count

Queue fill

index=_internal source=*metrics.log* group=queue host=<indexer> earliest=-4h
| eval fill_pct=round(current_size_kb/max_size_kb*100,1)
| timechart span=5m perc90(fill_pct) by name

To list only blocked queues:

index=_internal source=*metrics.log* group=queue blocked=true earliest=-4h
| stats count by host, name
| sort - count

Indexing lag

index=<index> earliest=-1h
| eval lag_seconds=_indextime-_time
| stats avg(lag_seconds) max(lag_seconds) by host, sourcetype

Parsing warnings

index=_internal sourcetype=splunkd earliest=-24h
  (component=DateParserVerbose OR component=LineBreakingProcessor OR component=AggregatorMiningProcessor)
| stats count by host, component
| sort - count

Forwarders connected

index=_internal source=*metrics.log* group=tcpin_connections earliest=-15m
| stats latest(_time) as last_seen sum(kb) as kb by hostname, version
| convert ctime(last_seen)

Dispatch count

| rest /services/search/jobs splunk_server=local count=0
| stats count sum(diskUsage) as bytes by label, author
| sort - count

Skipped searches

index=_internal sourcetype=scheduler status=skipped earliest=-24h
| stats count by savedsearch_name, reason
| sort - count

Memory by search

index=_introspection sourcetype=splunk_resource_usage component=PerProcess data.search_props.sid=* earliest=-4h
| stats max(data.mem_used) as peak_mem_mb latest(data.search_props.user) as user by data.search_props.sid
| sort - peak_mem_mb

License usage

Run on the license manager. Swap idx for st (sourcetype) or h (host) to change the split.

index=_internal source=*license_usage.log* type=Usage earliest=-7d@d
| eval GB=b/1024/1024/1024
| timechart span=1d sum(GB) by idx

Sources

HEC status codes and license thresholds were checked against the two Splunk pages below on the date at the top of this page. The remaining entries draw on general Splunk administration practice, and exact message wording and default values vary between versions, so confirm against the .spec files and documentation for your release before changing a limit.

You may also like...