When an AWS abuse report lands in your inbox about an EC2 instance running cryptocurrency mining malware, you know you’re in for an interesting investigation. This week, I conducted a complete forensic analysis of a compromised production server, and the findings highlight critical lessons about patch management, monitoring, and security hygiene.
Initial Discovery
The timeline started simple: AWS abuse report in December, 2025, about suspicious network activity. The client had already rebuilt a clean instance, but the compromised server remained running for forensic analysis. My task was to determine what happened, how it happened, and ensure the replacement was truly clean.
Forensic Snapshot Acquisition
First step: create an EBS snapshot without alerting any potential attackers still monitoring the system.
aws ec2 create-snapshot \
--volume-id vol-XXXXXXXX \
--description "Forensic snapshot for incident response" \
--region us-east-1
I exported the snapshot to a local disk image using qemu-nbd and mounted it read-only for analysis. This approach keeps the evidence pristine while allowing full filesystem access.
Malware Discovery
The bash history immediately revealed something suspicious: numerous SSH authentication failures in December, 2025, followed by cron job installations.
Four malware payloads were discovered:
- Cron-based downloaders executing every 60 seconds
- DoS malware running at 3:30 AM/PM daily from
/tmp/corn - AWS S3-hosted payloads from attacker-controlled buckets
- systemd persistence via
lived.serviceandalive.service
The attack pattern was clear: exploit, install persistence, download additional payloads.
Root Cause Analysis
Here’s where it got interesting. The instance was launched July, 2025, running OpenSSH 9.6p1-3ubuntu13.11. That version was vulnerable to CVE-2024-6387 (“RegreSSHion”), a critical unauthenticated remote code execution vulnerability patched in July 2024.
The instance ran for five months with a 13-month-old vulnerability. Why?
Checking the system logs revealed unattended-upgrades was configured and running daily. The package claimed success. But digging deeper:
# APT package list was never refreshed after deployment
ls -la /var/lib/apt/lists/
# Last update: September 2024 (deployment date)
# System thought it was up to date
apt list --upgradable
# No packages found
# Reality: 324 pending updates including critical SSH patch
The automated update system was broken at deployment. It ran daily for five months, successfully doing nothing.
Attack Timeline Reconstruction
- 13:00-14:00 UTC: 1,857 failed SSH attempts (exploitation pattern)
- 13:46:52 UTC: Malware cron jobs installed
- 20:22:55 UTC: Attacker SSH key added
Six hours between initial compromise and SSH key installation. The attacker used unauthenticated RCE, not stolen credentials.
- 05:01 UTC: Client team manually ran
apt full-upgrade - First successful
apt updatein instance lifetime - Discovered 324 pending packages
- Clean instance deployed
Database Extraction and Analysis
The application database required careful handling. I extracted a 100MB SQL dump and ran six security checks:
- Stored procedures (potential code execution)
- Stored functions
- Triggers
- File operations (
LOAD_FILE,INTO OUTFILE) - User-defined functions
- System commands
All checks came back clean. The database contained only data tables, no executable code. Safe for restoration.
Source Code Verification
The client needed their application code but didn’t trust files from the compromised server. The git repository was in Git, but the SSH key was passphrase-protected and the developer was unreachable.
I set up Docker isolation to clone and scan the repository without exposing my system:
docker run -it --rm -v ~/keys:/keys:ro alpine:latest sh
# Inside container: install git, clone repos, scan for malware
Security scans performed:
- ClamAV signature-based scanning: 2,036 files, 0 infections
- Webshell pattern matching: negative
- PHP backdoor detection: negative
- Obfuscated code search: negative
- Git history analysis: 4+ years, no suspicious patterns
The code was clean. The developer had been pushing from a local machine; the server just pulled updates.
Dependency Vulnerability Assessment
Here’s where the story gets worse. Using Grype to scan the Laravel application dependencies:
13 HIGH severity vulnerabilities
12 MEDIUM severity vulnerabilities
All have patches available
Checking git history for the last dependency update:
git log -1 --format="%ai %s" composer.lock
# 2025-05-23: Last update (7+ months ago)
The oldest HIGH severity vulnerability was published November, 2024. Multiple critical patches were released throughout November and December 2024, plus three fresh XSS vulnerabilities on January, 2025.
The developer hadn’t run composer update in over seven months.
Key Findings
Attack Vector: CVE-2024-6387 (RegreSSHion) - unauthenticated SSH RCE
Root Cause: Broken automated patching (deployed, enabled, non-functional)
Exposure: 150+ days with critical vulnerability
Data Breach: Assumed complete (619 user accounts, full client database)
Code Integrity: Verified clean via git source
Dependency Security: Severely neglected (13 HIGH CVEs unpatched for months)
Lessons for Infrastructure Teams
-
Test your automation. Enabled does not mean functional. The client’s automated patching was configured correctly but silently broken for five months.
-
Monitor your monitoring. No alerts fired when apt updates stopped working. No alerts when 1,857 SSH failures happened in one hour.
-
Security updates are not optional. A Laravel application handling sensitive data shouldn’t go 7+ months between dependency updates.
-
Patch immediately on critical CVEs. RegreSSHion was disclosed July 2024. This instance was deployed July 2025 already vulnerable. The deployment image was 10 months out of date before it even launched.
-
Isolate forensic analysis. Using Docker containers to scan potentially compromised code prevented any malware from touching my analysis system.
-
Verify, don’t trust. The database dump could have contained stored procedures executing malicious code. Always scan extracted data before restoration.
Remediation Steps
For the client, remediation required:
- Terminate compromised instance
- Deploy containerized application (AWS App Runner)
- Migrate to managed RDS for automated patching
- Force password reset for all 619 user accounts
- Rotate all credentials (DB, AWS, API keys)
- Update dependencies to patch 13 HIGH CVEs
- Implement dependency scanning in CI/CD
- Enable AWS GuardDuty for threat detection
- Configure CloudWatch alerts for failed authentication attempts
Conclusion
This investigation reinforced a fundamental truth: security tooling only works if it’s actually working. The client had configured automated updates, ran them daily, and assumed protection. Five months of false security ended with a compromised server, exposed customer data, and expensive incident response.
The silver lining? The source code was clean, the database was recoverable, and we caught the compromise before the attackers pivoted to other systems. Good monitoring would have detected this on day one. Better patch management would have prevented it entirely.
If you’re running infrastructure on AWS, check your instances right now. When did you last verify automated patching is actually applying updates? Not just running, but succeeding? That’s the gap that cost this client five months of exposure.