Blacksec

Administrator
Staff member
ROOT
VIP
Hey hackers - dark web dorks are not magic strings; they are the oldest query grammar in security, pointed this year at the clearnet seams where dark-web activity leaks out.
Run dark web dorks against paste mirrors, code hosts, and Trello boards - the network itself will not hand you a target.
TL;DR: Google has never indexed a single .onion page, so the working definition of a dark-web dork in 2026 is a clearnet query that catches residue of dark-web activity: leaked credentials, combo lists, forum dumps, onion addresses pasted in tickets.
The operator grammar - site:, inurl:, intitle:, filetype:, before:, -exclusions - is twenty-four years old and still load-bearing; `cache:` died in February 2024 and `related:` before it. Fifteen ready queries below, each one checked against current operator support. Neighbors: 40 more dorks, SQL dork set, combo lists.

What a dark web dork actually is​

The original Google Hacking Database was assembled by Johnny Long in 2002, migrated to Exploit-DB, and frozen in its growth years later - new operators stopped appearing around 2019, and the operators that died got replaced by nothing. Twenty-five or so operators remain in common use, and every one of them targets a search engine's index of the public web. That constraint is the whole design: a dark-web dork does not search the dark web; it searches everywhere the dark web has been quoted, mirrored, dumped, or indexed by mistake.
The leaks are the seams. Credentials harvested by infostealers land in combo lists; combo lists get pasted to Pastebin; Pastebin gets indexed. A market's onion address appears in a forum signature, a GitHub README, a support ticket, a Discord export. Seized forums get scraped and mirrored to clearnet archives. Every one of those events creates a queryable page, and the queries are ordinary dorks with the target swapped from a company name to a handle, a market, or a phrase unique to a leak.
The hosts rotate: pastebin.com draws the crawler attention, so leakers drift to paste.ee, dpaste, the German and Russian paste services, and single-use burn sites that never get indexed at all. A list frozen on one host stops finding residue the day the crowd moves; watching which paste domains show up in your last three runs' results is the cheapest way to notice the drift.
Detection runs both directions, which is why blue teams own this corpus too: the same query that finds your employee's password in a paste finds your company's own exposure. The difference between recon and monitoring is whose name sits in the query slot.
The lists themselves are public and old: GHDB mirrors on Exploit-DB, the aggregated GitHub collections, the reference pages that render the same operators behind a search box. What changes quarterly is not the grammar but the targets - this year's paste sites, this year's channel mirrors, this year's seized-domain redirects - so a dork file older than six months is missing the seams where fresh residue lands.

Operators that still work in 2026​

The grammar is small. Six operators cover nearly every dark-web dork worth running; the rest are refinements. The table is the cheat sheet - everything in the top-fifteen list below is assembled from these parts.
OperatorMatchesNotes
site:domain or path scopemust be first; scopes the rest
inurl:text in the address.git, .env, admin, login
intitle: / intext:page title / body textintitle pulls listings, intext pulls content
filetype: / ext:file extensionenv, sql, log, pdf, key
before: / after:index date rangedate-scoped to a breach window
-term, ORexclusion, unionnoise control, alternate phrasings
Two dead operators keep appearing in copied lists: `cache:` was removed by Google in February 2024, and `related:` went earlier - lists shipping either one are recycled from 2018 and should be treated as a freshness marker for the whole page they came from. AROUND(n) for near-matches still functions but rarely changes results in leak hunting, where quoted strings do the heavy lifting.
Placement rules decide whether a query works at all: site: goes first or it gets ignored, quoted phrases keep multi-word handles intact, and an unquoted passphrase becomes an AND of unrelated words that matches half the internet. Negation with -site: is the cheapest noise filter available - exclude your own domains before anything else, or your recon will report your own marketing pages back to you as findings.

Google has never seen an onion​

Every mainstream crawler skips .onion services: the addresses are not DNS, they change without notice, and hosting a crawler inside Tor at scale is its own operational problem. Bing, DuckDuckGo, and Google therefore return exactly zero onion results for any query, which means a "dark web dork" pasted into a normal search box is really hunting the clearnet's reflection of dark-web activity.
Engines that claim otherwise - "Google for the dark web" landing pages with a search box and a suspicious number of ads - are either running their own tiny crawler (nothing wrong with that) or running a phishing kit (very much wrong with that).
The genuine onion search engines - Ahmia, Torch, Haystak - maintain their own Tor-resident crawlers and their own index, so the operator subset they accept is smaller: quoted phrases, basic filtering, no Google grammar. Keep the grammars straight: full dark web dorks belong on the clearnet engines, and clean keyword search belongs on the onion indexes.

The top fifteen​

Assembled for exposure work: each dark web dork hunts credential residue, leak mirrors, or infrastructure traces with a placeholder company name. Swap TARGET and the queries run as-is.
#DorkCatches
1filetype:env intext:"DB_PASSWORD"stray database env files
2intitle:"index of" ".env"exposed directory listings
3site:github.com intext:"password" intext:"TARGET"creds committed to code
4site:github.com "BEGIN RSA PRIVATE KEY"private keys in repos
5filetype:sql "INSERT INTO" "password"dumped database files
6site:pastebin.com "TARGET" "password"paste-site credential slips
7site:pastebin.com "combo" OR "combo list"fresh combo lists at the source
8inurl:.git site:TARGET.comunprotected git metadata
9intitle:"index of" inurl:ftpopen file listings
10filetype:log intext:"username" intext:"password"leaked application logs
11site:trello.com "TARGET" "confidential"boards pasting internal docs
12intext:"@TARGET.com" "leaked" OR "breach"breach chatter naming the org
13intext:"onion" intext:"TARGET"your org's name near onion refs
14site:telegram.org "TARGET" "leak"public telegram post mirrors
15"TARGET" "dump" after:2025-01-01date-scoped dump mentions
Queries 6, 7, and 14 are the dark-web-specific ones: paste sites and public telegram channels are where combo lists surface first, hours before any monitoring platform picks the thread up. Query 13 inverts the usual direction - it does not hunt leaks, it hunts your organization's footprint on pages that talk about the network, which is how accidental onion exposure gets found.
Customization is where the list earns its keep. TARGET gets three passes - the brand, the brand without Inc, and the brand misspelled the way a leaker would type it at 3 a.m. - because combo lists are typed by tired humans, not exported from CRM. Non-English brands add the local-script variant; anything with a marketplace presence adds the vendor handle. Fifteen queries is the base grid; the twenty minutes spent adding your own variants is what makes it yours.

Reading a hit: from query to evidence​

A result is a lead, not a finding. The path from hit to evidence has three gates: authenticity, freshness, and scope. Authenticity first - a pasted credential pair proves nothing until you check whether the paste mirrors a known combo list or a staged honeypot, and a repo leak means nothing until the key is confirmed active. Freshness second: a 2019 paste of a password rotated quarterly is noise, and date operators exist to enforce that gate mechanically rather than from memory.
Scope third, and this is where recon becomes a report: one leaked pair is a finding, the same pair across four paste sites is a pattern, and the pattern plus a reachable service is an incident. The dork does not care about any of that - the dork only fetches candidates, and the analyst's job is the part no operator grammar can express - dark web dorks fetch, analysts judge. Forgotten subdomains follow the same logic: takeover-prone hosts and stale DNS records get found by query, then confirmed by hand.
The repeat runs matter more than the clever ones. A quarterly pass over queries 1, 3, 6, and 7 - the core dark web dorks for brand exposure - catches the slow leaks; a one-time marathon of every dork on the internet catches a screenshot. Build the list, schedule the list, keep the list short enough that every hit gets read.
Evidence handling is the unglamorous finish: capture the page text and the query that found it, timestamp both, and store them somewhere that is not the browser cache. Screenshots without the query string are anecdotes, and anecdotes do not survive a conversation with whoever has to act on the finding three weeks later.

The grammar is portable​

Dark web dorks written for Google mostly transfer to Bing and DuckDuckGo with one warning: operator support differs, and unsupported operators degrade silently - Bing returns results that ignore the part it did not understand instead of erroring. Bing's own grammar adds contains:, ip:, near:, and feed:, which Google lacks; DuckDuckGo accepts filetype:, site:, and inurl: reliably.
The habit that saves time: run the canonical list first on Google, then re-run the survivors on Bing, because Google's index and Bing's index disagree about paste sites more often than either admits.
Automation is the boring layer that makes any of this repeatable. Pagodo queries Google programmatically from a dork list, GooFuzz and the family of wrappers around it do the same with different flags, and the aggregated collections - OneDorkForAll's categories including its dark-web set, the bug-bounty dork repos, dorksearch-style reference pages - exist so nobody rebuilds the operator list from memory.
Treat any of these as a list source, not an oracle: the value is in the file you curate, prune, and re-run on a schedule - the difference between dark web dorks you own and dark web dorks you borrowed once.
Every scheduled run should be boring on purpose: same list, same order, same logging line per artifact. The moment a run becomes a hunt, the cadence breaks, and the cadence is the product - a list nobody re-runs is a bookmark, not a sensor.
Scale changes the failure mode. Ten hand-run queries per quarter barely register; a few thousand automated queries per day hit rate limits, CAPTCHAs, and the quiet indexing throttles that make automated results drift from what a human at a keyboard sees. Keep automation on your own targets and brand terms, keep the broad internet-wide sweeps manual or throttled, and log which query produced each artifact - an evidence trail with no query record is a screenshot with extra steps.
The second-order use is comparison: diff each run's result set against the last run. New hits are leads; vanished hits are leads too - a paste deleted after appearing in your results is an event worth writing down, and a repo that went private after a dork flagged it tells you somebody read the same page you did. The list is a sensor, and sensors are read by comparing readings, not by taking one.

The stance that holds​

Point dark web dorks at the clearnet seams - pastes, repos, boards, channel mirrors - keep the dead operators out of your lists, and automate only what you will still read next quarter. Fifteen queries run monthly against your own brand will surface more real exposure than a thousand scraped from somebody else's list and never executed. The grammar has not changed since 2002; the discipline of re-running it has never been optional.