This guide describes the USGenWeb Project as a genealogical source, explains why its content has historically been difficult to search, and provides details on how to use the GenWebSearch tool effectively, including how to interpret coverage ratings, spelling-variant disclosure, index freshness, and negative search results.
The USGenWeb Project as a Genealogical Source
What Is USGenWeb?
The USGenWeb Project is a volunteer effort, running continuously since 1996, in which local coordinators maintain free genealogy websites for individual U.S. counties. Across thousands of county sites, volunteers have transcribed an enormous body of source material: cemetery readings, vital records, census returns, deed abstracts, obituaries, biographies, court records, and local histories.
Much of this material exists nowhere else online. A volunteer who walked a rural cemetery in 1999 and typed every stone into a county page created a record that may now be more complete than the cemetery as it stands today. Deed abstracts and court minutes transcribed by county coordinators frequently cover record groups that the large commercial databases have not incorporated. USGenWeb at times may include data that is not available elsewhere, or research leads that lead to primary records.
Why USGenWeb Has Been Hard to Search
Despite its value, the USGenWeb corpus has three properties that make it challenging for typical research workflows:
It is fragmented across thousands of independently maintained sites. There is no central database. Each county site has its own structure, navigation, and hosting arrangement, reflecting thirty years of volunteer decisions. Searching “the USGenWeb Project” has really meant visiting county sites one at a time.
General-purpose search engines index it poorly. Much of the content lives in aging HTML with minimal structure, in long compiled pages, or inside PDF and Word documents. Search engines tend to deprioritize this kind of content, and even when a page is indexed, a surname buried in row 500 of a cemetery transcription rarely surfaces in general web results.
Much of the content is largely invisible to commercial genealogy platforms. A researcher who has “searched everywhere” may not have searched here. And if they had, it was likely one county at a time.
What GenWebSearch Provides
GenWebSearch is a full-text search across a purpose-built index of USGenWeb county sites, constructed from two crawl waves: a nationwide crawl in June 2026 and a targeted supplement in July 2026 that added several areas worth of previously missed county sites. The index contains:
- 388,651 pages of extracted text, including text extracted from PDF and Word documents found on county sites
- ~665 million words of searchable content
- 2,648 counties that yielded indexable pages, out of roughly 3,800 county-site crawls across the two waves
Measured against the roughly 3,145 counties and county-equivalents in the United States, that means about 84% of U.S. counties have at least some indexed content — though, as the coverage badges below make clear, “some” spans everything from a basic landing page to millions of words.
One search box now searches this entire corpus with records ranging the full spectrum. This is the tool’s core value: it converts a source you could only browse one county at a time into a source you can search.
Two points of orientation before you begin:
The index is a snapshot, not a live search. Results reflect what pages said when they were crawled. Sites change over time and some links will eventually die.
GenWebSearch is a finding aid, not a replacement. Every result links back to the original page, which remains the work of its USGenWeb volunteer transcriber. The tool exists to bring researchers to the county sites and not to substitute them. This website has no affiliation with USGenWeb, and the goal is simply to make these records easier to locate.
Consideration: A Hit Is Evidence; A Miss Is Not Evidence of Absence
A hit is straightforward: the indexed text contains your search terms, and you can click through to evaluate the page.
A miss is not straightforward in what it demonstrates. Of note, the coverage of the index varies enormously from county to county. Some counties have millions of words of dense transcription indexed; others yielded only a thin landing page. A county may also have little genealogical information on the site; however, we cannot rule out scenarios where GenWebSearch was unable to adequately crawl a site successfully or completely. In other words, GenWebSearch cannot claim full coverage of any county, let alone all of them.
To make the point concrete: the index is extraordinarily skewed. The top 1% of counties by size hold about 19% of all indexed words; the top 10% hold about 65%; and the bottom half of counties together hold only about 2%. A national word count of 665 million says nothing about whether your county of interest is one of the observed massive sites versus one of the stubs. That is what the “coverage badges” are for.
GenWebSearch is built around what we call Honest UX, in that we make every effort to ensure the interface tells you exactly what was searched, how each county is covered, and where the index falls short. The goal is to support users in their ability to interpret results. The features described below exist to serve that principle.
Coverage Badges
Every result card carries a coverage badge for its county, computed from a single metric: the total number of indexed words for that county. It is acknowledged these are qualitative terms applied for the sake of categorically parsing the corpus when viewing large sets of results.
- Substantial — 250,000 or more indexed words (currently 561 counties, about 21% of rated counties)
- Moderate — 25,000 to 250,000 indexed words (950 counties, about 36%)
- Limited — fewer than 25,000 indexed words (1,137 counties, about 43%)
Word count was chosen over page count deliberately since page count can be misleading given the large variation in how the county sites were structured.
Independently of its tier, a county may carry a partial crawl flag. Due to practical constraints, the crawler allotted each county a fixed time budget, and 259 counties had more content than the budget allowed. Their live sites contain more than what was indexed. The badge tooltip discloses this wherever it applies. Addressing these coverage gaps is on the roadmap for tool improvement.
Known Gaps in the Index
These are the specific, known ways the index falls short of the full USGenWeb corpus.
The RootsWeb loss — with a partial recovery. Many USGenWeb county sites historically hosted their actual record transcriptions on RootsWeb, a separate platform. During the crawls, RootsWeb-hosted pages refused access with no content being indexed. Of note, there were 24 counties whose RootsWeb-blocked content losses were confirmed, but with many found in July hosted on alternate sites: three Indiana counties (Delaware, Jasper, and Starke) were recovered in the July wave via replacement sites — Delaware County went from total loss to a Substantial rating. As of July 2026, nine counties remain without any indexed pages: Wapello (IA), Rowan (NC), Perry (OH), King William (VA), and five in South Carolina — Abbeville, Edgefield, Lancaster, Orangeburg, and Sumter. The remaining twelve confirmed-blocked counties are indexed only thinly. If your research touches these counties, go directly to their live USGenWeb sites to confirm directly what content resides there.
Time-capped counties. The 259 counties that hit the crawler’s time budget hold more on their live sites than the index shows. Again, this is on the roadmap for future improvement.
Certain pages the crawler could not parse or reach. A small number of state and county sites use structures (JavaScript-driven menus, image maps, unusual hosting, heavily templated pages) that defeated the crawler. Wyoming is the clearest example: 23 of its counties are technically rated, but the entire state yielded only 69 pages, most of that in a single county — Wyoming is effectively unindexed. In the July wave, 16 Georgia counties timed out (retry candidates for a future crawl), and 37 sites that answered the crawler nonetheless yielded no extractable content. Some counties therefore show little indexed content despite having active sites. Evaluations are on-going to manually verify these entries.
Oversized and image-only documents. PDFs over 45 MB were not downloaded/indexed, and scanned-image PDFs with no text layer yield little extractable text. PDFs hosted outside a county’s own domain (a common arrangement for some state archives) were also not fetched, and this affects hundreds of counties to some degree. County-history books and atlases that are widely available elsewhere (archive.org, Google Books) were deliberately deprioritized in favor of record transcriptions.
How Searching Works
Surname Spelling Variants
With “Search spelling variants” turned on, your surname is expanded to known historical spellings drawn from a curated dictionary of 1,368 surnames. This dictionary derives from the VariantChronicles research behind our companion tool for the Chronicling America newspaper collection. If the surname of interest is not in the dictionary, a deterministic set of spelling rules is tried instead (Mc↔Mac, i↔y, -son↔-sen, terminal -e, and similar patterns, capped at a total of 12 variants). If no spelling variant is found, then the search is done on the surname as entered.
First-Name Proximity
First names are optional and are searched near the surname, in either word order — “John Brubaker” and “Brubaker, John” both match. Three proximity settings are available:
Adjacent — the names appear within 2 tokens of each other. This tolerates middle initials (“John B. Brubaker”) and constructions like “John and Mary Brubaker.”
Near — the names appear within 15 tokens. This is a token-distance window, not a grammatical sentence boundary; it is best described as both words within a passage.
Same page — both names appear anywhere in the document, with no proximity requirement. Use this with care. Many USGenWeb pages are long compiled documents such as full county indexes, multi-family cemetery listings, and in the extreme, single obituary files running to hundreds of thousands of words — and on such pages the two names may be entirely unrelated. A “Same page” hit is a lead to evaluate, not a connection to record, and in some cases is no more valuable than searching by surname alone.
The tightening effect is dramatic in practice: in one test, a first-name-plus-surname pair matched over 1,200 pages at the Same-page setting, about 325 at Near, and about 150 at Adjacent. Users can consider whether to approach their search by “starting wide” or “starting narrow” via these options.
Other Uses of the First-Name Field
The search was intended for the general use case of searching by surname and an optional first name, but the crawled database is agnostic to how you use these search fields. You can use the surname / first name combination to search for Smith / Bible. Perhaps you have an ancestor of surname Wedig and family lore says he died by falling from a tree. It may be useful to enter the phrase “fell from a tree” as the surname along with entering Wedig as the optional first name with whatever adjacency rules you want to try. The short version is you can search for any word using these fields that you find useful, as the database itself has no knowledge of what a “name” is.
The Year Filter and “Mentions Years”
Date information on result cards reads, for example, “Mentions years 1832–1911” and this wording is intentional. It means that four-digit numbers within that range were found on the page. It does not mean the page is a record created in those years. A county biography page written in 1998 that mentions an 1832 birth will show that year.
The year filter works the same way. Critically, pages with no detectable year signal, about 11% of the index, are always included in filtered results. A page can be highly relevant to your 1840s research question without containing a single four-digit year.
One caution at the deep end of the range: apparent “years” before about 1600 are unreliable. Numbers in that range on transcription pages are frequently ledger page numbers, record numbers, or other non-date figures that the tool’s extractor cannot distinguish from years. Genuine 16th-century mentions do exist (chiefly on family-line pages tracing European origins), but treat any pre-1600 year signal as a prompt to read the page, not as evidence of a colonial-era record.
Record-Type Tags
Type tags (cemetery, vital, census, and so on) are heuristic hints detected from page content and URLs. They are useful for narrowing but are not a guarantee of what a page contains. Cemetery is by far the largest category at roughly 155,000 pages, followed by obituary, biography, military, vital, and census material. About 18% of indexed pages carry no tag at all; the type filter includes untagged pages by default, and there is an option to hide them if desired. If you filter by record type, remember that a relevant untagged page is hidden only if you chose to hide it.
Using GenWebSearch as a Research Method
A Suggested Workflow
Start broad. Search the surname alone, with variants on and no state filter, to see the national shape of the name in the index. The variant disclosure will show you which spellings were included.
Narrow deliberately. Add the state filter, then a first name at the Near setting, then record-type or year refinements only as result volume demands. Each refinement risks excluding relevant material; add them one at a time so you know what each one costs.
Check the coverage badge before interpreting anything. A hit in a Limited county is just as valid as a hit anywhere else, but a miss in a Limited county is close to meaningless, while a miss in a Substantial county without a partial-crawl flag is a genuinely informative negative result.
Log negative searches with their scope. Because variant lists are deterministic and every search discloses exactly what was searched, a GenWebSearch miss can be recorded in your research log as a properly scoped negative finding: the terms and variants searched, the filters applied, the county’s coverage badge and any partial-crawl flag, and the crawl vintage (June or July 2026). A negative search recorded this way meets the reasonably-exhaustive-search standard far better than “searched GenWebSearch, nothing found.”
Click through and evaluate the original page. The snippet shows what the page said when indexed. Read the full page in context: who transcribed it, from what source, and when. Note that a result page reflecting the 2026 snapshot may have since been updated on the live site.
Follow the derivative to its original. A USGenWeb transcription is a derivative source. It is often excellent, but for proof-argument purposes it points you toward the original record (the deed book, the stone, the courthouse register). Cite the transcription for what it is, and pursue the original where one survives.
When a Miss Should Send You Elsewhere
A miss in a Limited or partial-crawl county is an instruction, not a dead end. Visit the county’s live USGenWeb site directly and browse its record listings, as the content may exist but be unindexed. Check whether the county’s records were historically hosted on RootsWeb; if the county appears in the known-gaps list above, assume the index simply cannot see its content. And remember that the index deliberately deprioritized county-history books available on archive.org and Google Books; if your research question involves published county histories, search those platforms directly.
Interpreting Snippets and Handling Stale Links
USGenWeb sites are living volunteer projects: pages change, move, and occasionally disappear, and volunteers may add new information. The snippet on a result card shows what the page said when indexed; the live page may differ. When a link fails, try the county’s current USGenWeb site directly since the content is often still there under a new address. Because the index is a snapshot, newly added transcriptions will not appear in search results until a future re-crawl.
Relationship to the USGenWeb Project
GenWebSearch is an independent research tool with no affiliation with the USGenWeb Project or its state and county coordinators. The transcriptions it indexes are the work of thousands of volunteers over three decades; every result links to their pages, and researchers who benefit from a county site’s content are encouraged to note the transcriber’s credit line and to support the project’s volunteers and USGenWeb where opportunities exist.
Data Handling
GenWebSearch maintains a search index of text extracted during its crawl, used solely to identify which pages match a query. Results link to the original pages on USGenWeb sites; the tool does not republish county-site content as a destination in itself, and page content is not presented as a substitute for visiting the source. When you click through to a result, you are reading the volunteer’s page on their site.
About the Author
Nathan is an avocational genealogist and the founder of Evidence Toolbox. His research practice is grounded in the Genealogical Proof Standard, with primary-source work conducted at major repositories including the Library of Congress. He builds the tools on this site to solve problems encountered in his own research, and field-tests each one against real family lines before release. You can reach him at contact@evidencetoolbox.com.