External Leak Detection + Internal Sweep with WebSSH Login Validation
External Leak Detection + Internal Sweep with WebSSH Login Validation
A production module inside a Security Command Center (SCC) platform that hunts compromised credentials from two directions, outside-in (employee/customer credentials leaked to threat-intel feeds) and inside-out (internal hosts still running default/weak credentials), proves which ones actually work with autonomous validation, and turns them into a prioritized, analyst-ready queue. SSH validation during internal sweeps is performed through a built-in WebSSH terminal.
// Background
Leaked-credential alerts flood the SOC. Most are
noise, the dangerous ones need proof, by hand.
No signal-to-noise filter
Thousands of alerts: duplicates, dead URLs, and
malformed parses (url="https", 2-char passwords).
Most refer to already-rotated credentials
Manual validation won't scale
The only way to know a leak is live is to TRY the
login, open page, type creds, judge, screenshot.
Slow, inconsistent, impossible at volume
No consistent prioritization
Without a confirmed valid verdict, every alert
looks equally urgent, live leaks sit beside junk.
Effort lands everywhere except the real risk
No audit trail or feedback loop
Analyst judgments lived in spreadsheets and
inboxes, disconnected from the source data.
Same calls re-made, no way to correct safely
Multi-tenant blindness
Leaks must be attributed to the right business
unit / brand, raw feeds don't provide it.
No per-tenant ownership or routing
Net effect
Genuinely live leaks sit in the same pile as junk -
invisible until someone tries each one by hand.
// Solution
Data sources
Threat-Intel API
leaked-credentials feed
hourly ingest
Object Storage + Telegram
file uploads · combo lists
analyst + channel feeds
Pipeline
Ingest
Airflow
dedup · hourly
Ingest
Classify
BERT · extract model
multi-worker
Precheck
AI Agents
browser + LLM
try login · multi-step
Validate
Materialize
atomic swap
multi-worker
Fact
Portal
Vue 3 + FastAPI
+ feedback
Serve
↳ other classes → routed to their own flows (out of scope)
Airflow · multi-worker classify (BERT + extract) · browser agent · feedback at read-time · multi-tenant
// Impact
Validation: manual → autonomous
Before · by hand
hours · inconsistent · doesn't scaleAfter · browser agent
minutes · automated24/7
autonomous login testing at scale
Every verdict carries a login screenshot
+ a replayable agent trace.
Operational impact
Noise filtered pre-human
ML + SQL drop junk & dupes
Confident prioritization
work a confirmed-valid queue, not a pile
Explainable, evidence-backed
written reason + signals per verdict
Trustworthy feedback loop
overrides audited & reversible
Multi-tenant brand attribution
routed to the right business unit
Closed loop to remediation
valid leaks → notify + remediation log
Results
Units covered
3
Evidence-backed
100%
Noise filtered
~55%
Leak → validation
~3 min
Live volume & valid-rate tracked per
tenant & source in portal dashboards.
| Layer | Technologies |
|---|---|
| Orchestration | Apache Airflow (sensor-coupled ingest → process stages), scheduled notebook execution |
| Ingestion | Commercial threat-intel API, object storage (S3-compatible), Telegram, Python requests, pandas |
| ML classification & extraction | Two fine-tuned BERT models (HuggingFace Transformers), one classifies the alert, one extracts the URL / username / password, served via Ray Serve and loaded from the MLflow registry; DSPy + LLM available as an alternative |
| Active validation | External: browser-use autonomous agent driven by an LLM over a headless browser; Internal: per-protocol login sweep using WebSSH to attempt SSH logins; screenshots + run traces captured as evidence |
| Data store | PostgreSQL (raw landing + analytics fact tables), psycopg2 / SQLAlchemy, pooled connections |
| Backend API | FastAPI, psycopg 3, Pydantic, in-memory cache, async workers |
| Frontend | Vue 3, Quasar, Vite, TypeScript, Pinia, TanStack Vue Query, ECharts (incl. Sankey) |
| Analyst workflow | Automated analyst notifications, non-destructive feedback overlay (valid / invalid / hide / remediated) |
| Secrets / infra | Infisical (secret management), nginx → Uvicorn/Gunicorn workers, Docker |
Leaked credentials arrive from external feeds, internal systems are swept for default and weak passwords, and an AI agent validates each candidate before it reaches the portal.
Disclaimer: sample visuals may contain anonymized, simulated, or non-production values for presentation purposes.