GenWebSearch
One search across the full text of USGenWeb's volunteer-transcribed county sites — cemetery transcriptions, vital records, biographies, obituaries, court records, and more. A hit is evidence; a miss is not evidence of absence — every search tells you exactly what was searched and how well each county is covered.
Why USGenWeb Is Hard to Search — and Why It’s Worth It
The USGenWeb Project is one of the oldest volunteer efforts on the genealogical web: since 1996, county-level volunteers have been transcribing and publishing cemetery readings, vital records, will abstracts, biographies, obituaries, tax lists, and court records — much of it material that exists nowhere else online, transcribed from courthouse books and cemetery walks by people who live where the records are.
The catch is structure. USGenWeb is not one site; it is thousands of independent county sites, built by different volunteers across three decades, with no shared search. General web search engines index this material unevenly — older, plain-HTML transcription pages are exactly the kind of content modern search tends to bury — so researchers are usually left clicking county by county, page by page, hoping.
GenWebSearch replaces that with a single full-text search across 388,000 indexed pages from USGenWeb county sites. Enter a surname and the tool searches every indexed page — optionally expanding to historical spelling variants — and returns results with the source county, a snippet of the matching text, and a direct link to the volunteer’s original page, where the transcription and its source citation live.
To learn more about GenWebSearch — what the index contains, how coverage badges work, and how to interpret results — read the GenWebSearch Research Guide.
How to Run a Good Search
- Start with the surname alone, variants on. The variant expansion draws on a curated dictionary of historical spellings and falls back to spelling rules for names it doesn’t know — the same principle that finds a Brubaker filed under Brubacher or Brubaker.
- Read the coverage disclosure before you trust a miss. Every result set states exactly what was searched, and county coverage varies enormously — some counties have hundreds of transcribed pages, others almost none. A hit is evidence. A miss in a thinly covered county is close to no information at all; a miss in a deeply covered one is worth documenting as a negative search.
Frequently Asked Questions
About the Tool
What is GenWebSearch?
GenWebSearch is a free full-text search across a purpose-built index of USGenWeb Project county sites. The USGenWeb Project is a volunteer effort, running since 1996, in which local coordinators maintain free genealogy websites for individual U.S. counties spanning cemetery readings, vital records, census returns, deed abstracts, obituaries, biographies, court records, and local histories, much of it available nowhere else online. GenWebSearch lets you search this material with one search box instead of visiting county sites one at a time.
Is GenWebSearch affiliated with the USGenWeb Project?
No. GenWebSearch is an independent research tool with no affiliation with the USGenWeb Project or its state and county coordinators. The transcriptions it indexes are the work of thousands of volunteers over three decades. Every result links back to the volunteer’s original page, and the tool exists to bring researchers to the county sites, not to substitute for them.
How big is the index?
The index contains 388,651 pages of extracted text — roughly 665 million words — drawn from 2,648 counties, out of roughly 3,800 county-site crawls. Measured against the roughly 3,145 counties and county-equivalents in the United States, about 84% of U.S. counties have at least some indexed content. Coverage is very uneven, however: the top 10% of counties hold about 65% of all indexed words, and the bottom half together hold only about 2%. That is why every result carries a per-county coverage badge.
When was the content crawled? Is this a live search?
The index is a snapshot, not a live search, built in two crawl waves: a nationwide crawl in June 2026 followed by a supplemental crawl in July 2026. Results reflect what pages said when they were crawled. Content added to county sites since then will not appear until a future re-crawl.
Why isn’t this content on Ancestry?
USGenWeb transcriptions are not part of any commercial or aggregated genealogy index. The corpus is fragmented across thousands of independently maintained volunteer sites, much of it in aging HTML or inside PDF and Word documents that general search engines index poorly. A researcher who has “searched everywhere” often has not searched here, which is exactly the gap this tool fills.
Searching
How do spelling variants work?
With “Search spelling variants” turned on, your surname is expanded using a curated dictionary of 1,368 surnames covering genuine historical variants — Smith/Smyth, Mc/Mac prefixes, -son/-sen endings, anglicized forms, and the like. If your surname is not in the dictionary, a deterministic set of spelling rules is tried instead (Mc↔Mac, i↔y, -son↔-sen, terminal -e, and similar, capped at 12 variants).
What do the Adjacent, Near, and Same page settings mean?
These control how close a first name must be to the surname. Adjacent requires the names within 2 tokens of each other, which tolerates middle initials and constructions like “John and Mary Brubaker.” Near requires them within 15 tokens — a token-distance window, not a sentence boundary. Same page requires only that both names appear somewhere in the document. Use Same page with care: many USGenWeb pages are long compiled documents (full county indexes, multi-family cemetery listings), and on such pages the two names may be entirely unrelated. A Same-page hit is a lead to evaluate, not a connection to record. The effect is large in practice — a name pair can match over a thousand pages at Same page but only a few hundred at Near and under two hundred at Adjacent.
What does “Mentions years 1832–1911” on a result card mean?
Exactly that: four-digit years in that range appear somewhere in the page text. It does not mean the page is a record created in those years — a county biography written in 1998 that mentions an 1832 birth will show that year. The year filter works the same way, and pages with no detectable year signal (about 11% of the index) are always included in filtered results, because a page can be relevant to your 1840s question without containing a single four-digit year.
A result claims to mention years in the 1500s. Is that real?
Treat any year signal before about 1600 as unreliable. Numbers in that range on transcription pages are frequently ledger page numbers, record numbers, or other non-date figures that the extractor cannot distinguish from years.
How reliable are the record-type tags (cemetery, vital, census…)?
They are heuristic hints detected from page content and URLs, and not a guarantee of what a page contains. About 18% of indexed pages carry no tag at all. The type filter includes untagged pages by default and you may choose to hide them if you wish.
Coverage and Negative Results
What do the coverage badges mean?
Each result carries a badge for its county, based on total indexed words: Substantial (250,000+ words), Moderate (25,000–250,000), or Limited (under 25,000). These labels are a way to subjectively categorize the search results by county size.
I searched a name and got zero results. Does that mean the records don’t exist on USGenWeb?
No. Some sites were only partially crawled, and the tool’s crawl may have missed records.
What is the “partial crawl” flag?
The crawler gave each county a fixed time budget, and 259 counties had more content than the budget allowed — their live sites contain more than what was indexed. The badge tooltip discloses this wherever it applies, independently of the county’s tier. A miss in a partial-crawl county should send you to the live site, not into your research log as a negative finding.
What content is known to be missing from the index?
The known gaps include: RootsWeb-hosted content (many county sites historically hosted their transcriptions on RootsWeb, which blocks automated access — nine confirmed-affected counties have no indexed pages at all, five of them in South Carolina, and Hawaii’s five county sites remain fully blocked); 259 time-capped counties; sites the crawler could not parse (Wyoming’s templated sites are the clearest case — 23 rated counties but only 69 indexed pages statewide); and oversized or image-only documents (PDFs over 45 MB were not downloaded, scanned PDFs with no text layer yield little text, and county-history books widely available on archive.org and Google Books were deliberately deprioritized in favor of record transcriptions).
Results and Links
A result link is dead or the page looks different from the snippet. Why?
USGenWeb sites are living volunteer projects: pages change, move, and occasionally disappear. The snippet shows what the page said when it was indexed; the live page may differ. When a link fails, go to the county’s current USGenWeb site directly — the content is often still there under a new address. Which crawl wave your county came from also matters: most states reflect June 2026, while the supplement states (Arkansas, Florida, Georgia, Kansas, Louisiana, Nebraska, Nevada, South Dakota, West Virginia, Wisconsin, DC, and most of Indiana and New York) reflect July 2026.
Is a USGenWeb transcription good enough to cite as a source?
Cite it for what it is: a derivative source. Volunteer transcribers tend to be careful and document their work, and for counties with courthouse fires or thin surviving records a transcription is sometimes the only accessible trace of a destroyed original. But for proof-argument purposes it points you toward the original record.
Does GenWebSearch republish the county sites’ content?
No. The index stores extracted text solely to identify which pages match a query. Results link to the original pages on USGenWeb sites, and page content is never presented as a substitute for visiting the source. When you click through, you are reading the volunteer’s page on their site — and researchers who benefit from a county site’s content are encouraged to note the transcriber’s credit line and support the project’s volunteers where opportunities exist.
Related Reading
- Negative Searches: Why “Not Finding It” Still Matters — this tool tells you when a miss is meaningful; that article explains what to do with it.
- VariantChronicles — the same variant-search philosophy applied to 23 million newspaper pages in Chronicling America.
- AbsenceOfRecord — turn a documented miss into a standards-ready negative search statement.
Stay Updated
New tools ship regularly. Subscribe for occasional updates on new genealogy research tools — no more than once or twice a month.
SubscribeEvidence Toolbox is free and independent. Support it on Ko-Fi.