Abstract
This study documents an original, controlled experiment and its accompanying analysis implementation. A Linux virtual machine generated real OpenSSH records while an operator script performed successful logins, an isolated incorrect password, concentrated failures against an existing account, and failures against a nonexistent account. A Python tool preserves record provenance, reconstructs a UTC timeline and applies transparent sliding-window review rules.
The penalty-disabled control acquisition contains 11 primary failed-authentication records and 3 accepted-authentication records. The independent action ledger agrees with these counts. Ancillary invalid-user notices are preserved without being counted as additional password attempts. The experiment measures behavior in this small, labeled corpus; it does not estimate detection accuracy on Internet traffic or establish that an account was compromised.
Research question and contribution
Can a reproducible SSH authentication timeline distinguish an isolated mistake from concentrated failure patterns without overstating compromise? The contribution is a complete, inspectable experimental chain: declared scenarios, real daemon records, an independent client ledger, a source-preserving parser, parameterized detection, threshold sensitivity and downloadable evidence.
The research does not claim a new SSH vulnerability, a CVE, or a novel invention of log analysis. Its original contribution is this experiment, the generated dataset and the analysis implementation. Authenticated users can make mistakes; failures and later success therefore trigger review rather than an automatic conclusion about an attacker.
Methodology and laboratory isolation
The laboratory used Alpine Linux 3.24.2, Linux 6.18.52-0-virt, OpenSSH 10.3p1 / OpenSSL 3.5.8, VirtualBox 7.2.20r175154. It booted from a SHA-256-verified official Alpine image, with a documented local boot-console adjustment. The VM had no attached hard disk and no enabled network adapter; the serial console provided control and evidence export. The production website was used only as a publication destination.
The dedicated daemon listened on 127.0.0.1:2222. Password authentication was enabled solely to create controlled authentication records. Root login was disabled, the clients targeted only purpose-created laboratory identities, and the client and server limited each connection to one password attempt. The released configuration records the actual effective settings; these are experimental choices, not a production configuration recommendation.
During laboratory bring-up, an initial pilot failed at the transport stage: the RAM-only Alpine guest had no configured loopback address, and the client reported Network unreachable. Those transport errors were not counted as authentication failures; the raw pilot server corpus was not exported. The loopback address was configured in the runner and the experiment was restarted in a fresh guest. The authentic pilot client/interface capture and exclusion notes are included separately in the downloadable bundle.
Two executions are retained separately: one with the installed native penalties, and one with PerSourcePenalties no confined to this offline control. The scenario plan and client ledger are independent of the parser. Successful connections executed only a harmless marker command. In the control run, each failure used one deliberately incorrect password; there was no credential search. The isolated mistake was separated from the concentrated scenarios by more than the detector's 60-second window. The burst was followed by an explicitly legitimate success, allowing the review rule to be tested without labeling success as compromise.
| Scenario | Ground-truth purpose | Failed | Accepted |
|---|---|---|---|
| Legitimate baseline | Authorized successful logins | 0 | 2 |
| Isolated benign typo | One known incorrect password | 1 | 0 |
| Existing-account burst | Controlled repeated incorrect passwords | 6 | 0 |
| Nonexistent-account burst | Controlled invalid-account attempts | 4 | 0 |
| Legitimate login after burst | Correct credential supplied by authorized operator | 0 | 1 |
Timestamp handling. The guest used UTC and inherited the VirtualBox host wall clock. No network clock synchronization was possible in the offline VM. Guest environment/timing records and host submission/export UTC times are preserved. Clock accuracy and drift were not independently calibrated; BusyBox syslog timestamps have one-second precision. The parser requires an explicit year and timezone for BSD syslog records, preserves the original text, and emits RFC 3339 UTC timestamps. Second-level precision is retained; equal timestamps do not establish a causal order between separate connections.
Evidence preservation and provenance
Collection, examination, analysis and reporting follow the process described in NIST SP 800-86. Original daemon bytes are retained in sshd-syslog.log; the action ledger, version information, daemon configuration, experiment script and acquisition notes are included in the release. Derived reports are produced from a copy of the log.
The original sshd-syslog.log is 9870 bytes. Its SHA-256 digest is:
b6d85e3d58d5cf626d30a8195a23ba50ea466ce4cd186b4a93ce60c016fa2d07The evidence hash manifest lets readers check consistency of each released evidence file. A hash does not independently prove who generated a log, that the log is complete, or that its clock was accurate. This is a reproducible research acquisition with documented provenance, rather than a claim of independently witnessed legal chain of custody.
The public dataset contains only laboratory identifiers and loopback addresses. Credentials, private keys, hosting details and unrelated local files are outside the release. The machine-readable report carries the input hash, parser parameters and coverage counts; timeline.csv retains source line references.
Original sshd-syslog.log line 16 · Isolated benign typo; later failures are separated by 65 seconds.
Oct 10 19:15:21 cybery-ssh-lab auth.info sshd-session[1825]: Failed password for labuser from 127.0.0.1 port 40324 ssh2Original sshd-syslog.log line 44 · Supporting invalid-user notice; not an extra authentication failure.
Oct 10 19:16:26 cybery-ssh-lab auth.info sshd-session[1875]: Invalid user invaliduser from 127.0.0.1 port 51076Original sshd-syslog.log line 68 · Explicitly legitimate success after controlled failures; review indicator, not compromise.
Oct 10 19:16:26 cybery-ssh-lab auth.info sshd-session[1903]: Accepted password for labuser from 127.0.0.1 port 51114 ssh2Verified results: penalties-disabled control
- All 14 control-trial actions match 14 primary daemon records by account, outcome, loopback source and native timestamp within the client action interval. Successful clients returned 0; failed clients returned 255.
- The control contains 11 failed passwords: one isolated benign typo, six failures against labuser, and four failures against invaliduser. Four accompanying invalid-user notices do not add four more failed authentications.
- The isolated typo is 65 seconds before the first concentrated failure. The default five-failures/60-second detector produces 6 overlapping failure endpoints and 1 success-after-failures review candidate; the later success is explicitly legitimate.
- The separate installed-default trial records 2 accepted passwords, 5 failed passwords and 7 native connection refusals, despite 12 client status-255 results. The final planned legitimate login was refused before password authentication.
- Original evidence hashes were verified after console export, and an independent verifier agrees with the analyzer. The small controlled dataset does not demonstrate real-world detection accuracy or account compromise.
The parser recognized 35 event-bearing lines out of 74 input lines. It reports 37 SSH-related lines outside its recognized event patterns and 2 unrelated lines. Coverage is disclosed so that omitted message types are visible; it is not a recall score. The raw log remains available for manual inspection.
Native penalties change what reaches authentication
A separate execution retained the installed PerSourcePenalties defaults. Its ledger contains 14 client actions, but the daemon recorded only 2 accepted passwords and 5 failed passwords, followed by 7 connections refused by native source penalties. The final planned legitimate login returned client status 255 and never produced an accepted-authentication record.
The daemon explicitly logged a penalty activation, then connection drops with the reason penalty: failed authentication. The effective configuration records authfail:5, min:15 and max:600. The default trial reached four concentrated authentication failures after the isolated typo; accumulated native penalties then refused subsequent connections. The source-penalty records are policy context, rather than additional failed passwords.
| Observed measure | Installed default penalties | Control: penalties disabled |
|---|---|---|
| Client actions | 14 | 14 |
| Client status 255 | 12 | 11 |
| Primary failed-password records | 5 | 11 |
| Accepted-password records | 2 | 3 |
| Native penalty connection refusals | 7 | 0 |
| Five-failures/60s detector endpoints | 0 | 6 |
Default-policy trial · original log line 37
Oct 10 19:12:25 cybery-ssh-lab auth.info sshd[1792]: srclimit_penalise: 127.0.0.1/32: activating ipv4 penalty of 19.813 seconds for penalty: failed authenticationDefault-policy trial · original log line 38
Oct 10 19:12:25 cybery-ssh-lab auth.info sshd[1792]: drop connection #0 from [127.0.0.1]:52290 on [127.0.0.1]:2222 penalty: failed authenticationForensic implication: a nonzero SSH client result is not evidence of a failed password. Transport errors and native connection refusals must remain separate from authentication outcomes. In this default-policy capture, the five-failure rule did not reach its threshold even though the ledger records a concentrated sequence of connection attempts: native enforcement changed the available authentication records. This is an observed interaction between two controls in this corpus, not a general estimate of attacker activity or detector effectiveness.
OpenSSH documents the source-penalty mechanism; the OpenSSH 9.8 release notes describe its introduction. The installed build and actual effective settings, rather than current documentation alone, establish this experiment's configuration. The subsequent control run explicitly used PerSourcePenalties no to let the complete planned sequence reach authentication. That change is confined to this offline experiment and is not a production recommendation.
The default-trial raw log, action ledger, separate JSON analysis, client transcript, acquisition metadata and hashes are all released. Refusal records contain no account name; the deliberately sequential, single-source ledger and timestamps support the comparison, while same-second log order alone would not attribute concurrent activity on a real host.
Detection rules and threshold sensitivity
The default rule flags a source when at least five primary failed authentications fall inside an inclusive 60-second window. Source grouping uses the observed network address; all connections here share loopback, so the result describes one local source. A separate review rule flags an accepted authentication when the same source and account have at least five earlier failures in the previous 60 seconds.
The tool records evidence event IDs for every qualifying endpoint. Overlapping windows can produce several endpoints for the same burst. Endpoint counts are not counts of distinct attacks or incidents. Success after failures is a review indicator: this experiment deliberately creates it with a legitimate login.
| Failure threshold | Qualifying failure endpoints | Successes requiring review |
|---|---|---|
| 3 | 8 | 1 |
| 5 | 6 | 1 |
| 8 | 3 | 0 |
Changing the threshold changes what the rule flags, without changing the underlying evidence. Lower thresholds also flag more ordinary errors in some environments. The benign isolated mistake in this experiment does not meet the default threshold. This single small capture cannot provide generalizable precision, recall or false-positive estimates.
The controlled repeated-password behavior is consistent with the behavioral category described by MITRE ATT&CK T1110.001, Password Guessing. This mapping describes the experiment's behavior and does not attribute activity to a threat actor. No password was discovered by guessing.
Timeline reconstruction and interpretation
The normalized control capture spans 2026-10-10T19:15:20Z to 2026-10-10T19:16:26Z. Each recognized record has an event ID, source line number, original text, timestamp origin, event kind and available username, process ID, address and source port. Records are sorted by normalized time with original line order retained for ties.
An Invalid user notice and a following Failed password for invalid user can describe the same client connection. The parser counts the latter as the primary failure, while retaining the former as supporting context. Joining by address, source port and daemon PID can help review a connection, but PID reuse and missing fields limit that correlation.
An Accepted password record demonstrates that the daemon accepted that authentication method. It does not identify the human controlling the client or describe subsequent impact. The independent client transcript records the harmless command executed in this experiment; authentication logs alone would not prove command execution on an arbitrary investigated host.
Readers can inspect the complete CSV timeline, structured events and evidence references in the report, alongside the original log and action ledger. Failed client exit codes corroborate the controlled scenarios; they are not substituted for daemon evidence.
Limitations and threats to validity
- The datasets come from a single VM configuration, one OpenSSH build and two small, deliberately chosen sequences and one excluded setup pilot. It lacks background traffic, distributed sources, NAT, real attackers and production authentication policies.
- All attempts originate from loopback. No inference about attacker geography, source reputation or distributed password spraying is supported.
- Thresholds are transparent experimental heuristics, rather than calibrated classifiers. Many repeated failures can arise from a broken script, stale credentials or human mistakes.
- OpenSSH versions, PAM integration, distro logging, localization, rotated logs and journald exports may change the available records. Unknown lines are disclosed; the tool is not a universal SSH parser.
- BSD syslog lacks an explicit year and UTC offset. The year and timezone supplied for this acquisition are external context. Cross-year or mixed-timezone logs require separate handling; the tool does not silently infer missing chronology.
- Timestamp precision and the documented VM clock-setting procedure limit timeline confidence. Logs can be edited, lost or selectively recorded; integrity manifests do not establish completeness or independent authenticity.
- SSH authentication logs do not reveal all session commands, persistence, lateral movement or data access. This study does not claim a complete endpoint investigation.
- The native-policy trial required a documented manual finalize step after fail-fast stopped on the refused legitimate login. Each trial was preserved separately; there are no repeated-run confidence intervals.
- The article and code were prepared with AI assistance under the named author's publication workflow. Results come from executed laboratory activity and supplied evidence, with no claim of external peer review or program endorsement.
Security analysis and practical mitigations
For a production service, evaluate key-based authentication or certificates, disable password authentication where operationally feasible, prohibit direct root login, restrict permitted users, and limit network exposure through a VPN or an administrative allowlist. Where passwords remain necessary, consider a deliberately configured second factor through AuthenticationMethods. Changes must be validated against the installed OpenSSH version and tested without locking out administrators.
MaxAuthTries limits attempts within a connection; it does not impose a global rate limit across new connections. Complement it with appropriate connection controls and a monitored rate-limiting policy. Automated blocking must account for shared IP addresses and maintain a recovery path. A loopback laboratory threshold should not be copied blindly into production.
Preserve authentication logs centrally, use consistent UTC clock synchronization, document clock drift, define retention, and monitor collection gaps. Keep access to logs limited and preserve original evidence before normalization. LogLevel VERBOSE can support authentication investigation; DEBUG logging can disclose sensitive information and should be carefully scoped. Correlate authentication alerts with authorized changes, session activity and endpoint telemetry before drawing conclusions.
Reproduce the experiment and inspect the source
The reproducibility bundle includes the real evidence, analysis outputs, Python source and tests, laboratory runner, acquisition metadata and instructions. The lab has no Internet-facing endpoint. The source code can also be downloaded directly as ssh_forensic.py.
- Read the included README, acquisition notes and hashes. Use a disposable Linux VM with no network adapters, with the documented OpenSSH packages available offline.
- Run the included laboratory script as root inside that disposable VM. It creates only its dedicated lab account, temporary credentials and a dedicated loopback daemon, performs the declared scenarios and captures real logs. Never use it on a production system.
- Copy the public evidence files out through the console or an equivalent offline mechanism, excluding passwords and private keys. Preserve the raw bytes and record the acquisition context.
- Run the Python analyzer on the supplied capture, or on a new capture with its actual year and timezone. Re-running the experiment will change timestamps, PIDs, ports and hashes; agreement should be evaluated by scenario and record semantics.
python -m unittest discover -s tool -p test_ssh_forensic.py -v
python tool/ssh_forensic.py \
--input evidence/sshd-syslog.log \
--output analysis-reproduced \
--year 2026 --timezone UTC \
--threshold 5 --window 60Python 3.10 or newer is sufficient for the standard-library analyzer. report.json, timeline.csv, the Markdown report and SVG figures are derived artifacts. The included README explains the ledger comparison and hash verification commands. New analyses should preserve their own inputs and acquisition metadata.
References
- OpenSSH: sshd(8) — daemon execution, logging and configuration validation.
- OpenSSH: sshd_config(5) — authentication and logging controls; consult the installed version for exact availability.
- NIST SP 800-86: Guide to Integrating Forensic Techniques into Incident Response — evidence collection, examination, analysis and reporting.
- IETF RFC 3339: Date and Time on the Internet — normalized timestamps.
- MITRE ATT&CK T1110.001: Password Guessing — behavioral mapping for controlled repeated failures.
References checked on 2026-10-10. This research has not been approved or endorsed by Anthropic. Publication is evidence of research practice, not a statement of eligibility or acceptance into any verification program.
Author
Cybery di Mariani Yuri