This guide explains how to use BookTrace to search genealogy-relevant books held by the Internet Archive. It covers the Internet Archive’s significant genealogical holdings, why the Archive’s native search falls short for genealogical research, how BookTrace’s two-layer index works, how to interpret every element of a BookTrace result, and what the tool can and cannot tell you. The guide is written for intermediate and advanced researchers.
The Internet Archive as a Genealogical Resource
What the Internet Archive Holds
The Internet Archive (archive.org) hosts one of the largest freely accessible collections of digitized genealogical source material in existence. Among its tens of millions of digitized items are hundreds of thousands of genealogy-relevant books: county and local histories, published family genealogies, city and county directories, vital records abstracts, biographical encyclopedias, military records, passenger lists, and genealogical periodicals. Nearly all of this material is in the public domain.
The depth of the collection reflects its contributors. Digitized books have been supplied by the Allen County Public Library Genealogy Center, widely regarded as the premier genealogy library in the United States, along with Harvard University, the Library of Congress, the New York Public Library, the University of Michigan, Boston Public Library, Brigham Young University, the University of Toronto, the National Library of Scotland, and dozens of genealogical and historical societies. For many county histories and privately published family genealogies, the Internet Archive scan is the only digitized copy available anywhere, and in some cases the physical original survives in only a handful of libraries.
Why These Sources Matter for GPS-Standard Research
The Genealogical Proof Standard requires a reasonably exhaustive search. County histories, published genealogies, city directories, and vital records abstracts are exactly the categories of derivative and authored sources that Evidence Explained devotes major sections to. These are sources that serious researchers must consult and, when appropriate, must document having searched even when nothing is found. A researcher who has not examined the relevant county histories and published genealogies for a research subject has not conducted a reasonably exhaustive search.
The Internet Archive makes tens of thousands of these books available at no cost. The problem has never been access. The problem is findability.
The Weaknesses of Native Internet Archive Search for Genealogy
The Internet Archive’s search interface is built for a general audience.
- No genealogical search model. There is no way to search by surname as a surname, filter by geography as a research locality, or restrict results to a record type such as county histories or vital records abstracts. A search for “Hanks Kentucky” returns overwhelming noise — books about lakes, books by authors named Hanks, and books that merely mention Kentucky in a subject heading — with no way to separate the signal.
- Metadata search and full-text search are conflated. The default search runs against catalogue metadata (title, subject, description). Full-text search exists, but its coverage is opaque: a researcher cannot easily determine which books do and do not have searchable text, so a zero-result search is uninterpretable. Was the name absent, or was the book simply never text-searchable?
- No page-level pointers across the collection. There is no mechanism to ask, across thousands of books at once, “in which books, and on approximately which pages, does this surname appear?” Finding a single mention of a research subject buried on page 452 of a county history requires opening that specific book and searching it individually, assuming its text layer exists and its OCR is legible.
- No support for negative-search documentation. GPS-compliant research requires documenting what was searched and not found. The native interface offers no statement of scope, no record of what a search covered, and no citation-ready description of a null result.
- Uneven OCR quality with no clear disclosure. Much of the genealogical corpus was scanned from microfilm or consists of handwritten material. OCR quality ranges from excellent to unusable, and the native search gives no indication of which category a given book falls into.
In short: the Internet Archive is a magnificent library that has never had a card catalog built for genealogists. BookTrace is that card catalog.
What BookTrace Is
BookTrace is a free finding aid over 15,442 books catalogued from the Internet Archive’s genealogy and americana collections. It points researchers at specific books where a searched surname or place appears, with approximate page numbers, and links every result back to the original item on archive.org. BookTrace hosts no book content of its own. It is an index, a card catalog, over a library that already exists.
Two points of positioning matter from the outset:
- BookTrace is not a full-text search of the Internet Archive. It is a filtered, curated slice of the Archive built specifically for genealogists. A negative result in BookTrace is a statement about the 15,442 books in this index, not about the Archive’s 40+ million items.
- BookTrace is built around clear disclosure. Every design decision in the tool exists to tell you precisely what was searched, what was found, and what the tool could not see. The nuance on page numbers, the three-way split of results, the coverage statement attached to every search, and the negative-search citation language all serve that purpose. This guide explains each of those disclosures.
BookTrace is an independent product of Glenside Digital LLC, part of the Evidence Toolbox suite. It is not affiliated with or endorsed by the Internet Archive.
What BookTrace Can and Cannot Do
BookTrace can:
- Search 39.3 million extracted names and places across 7,179 fully text-indexed genealogy-relevant books in a single query
- Search the catalogue metadata of all 15,442 books in the index, including the majority whose text could not be read
- Expand surnames to spelling and OCR variants, given names to period abbreviations, and places to state-abbreviation equivalents, and report every expansion it applied
- Point you to approximate page numbers, saving hours of book-by-book manual searching
- Filter by probable source type, geography, and date, with explicit disclosure whenever a filter excludes unclassifiable or undated books
- Support multi-surname FAN and family-group searches with same-page and near-page proximity constraints
- Generate citation-ready documentation for every search, including BCG-appropriate negative-search language
BookTrace cannot:
- Search the text of the 8,263 books without usable OCR (no tool can, until they are re-imaged or transcribed)
- Show you the sentence or passage where a match occurred; it can only point you to the approximate page
- Confirm that two terms on the same estimated page appeared in the same sentence or refer to one another
- Reliably find women recorded in sources only under their husbands’ names
- Guarantee exact page numbers; every page reference is an estimate, marked with a tilde for that reason
- Stand in for a search of the entire Internet Archive, or of any repository beyond its own 15,442-book scope
- Replace reading the source. It is a finding aid.
When to Use BookTrace vs. Internet Archive Native Search
Both tools have real value for genealogical research; they solve different problems, and the most effective workflow uses each for what it was built for.
Reach for Internet Archive native search when:
- You already know which book you want to search and want to see the passage. IA’s “Search inside” returns exact page-level hits with the sentence around each match, which is faster than opening the book and scrolling.
- You are browsing the collection or exploring adjacent material. IA’s advanced search, collection filters, and subject headings are built for open-ended discovery.
- The book you’re looking for may sit outside genealogy or americana. BookTrace’s 15,442-book scope is deliberate, but IA holds tens of millions of items.
Reach for BookTrace when:
- You need to search across many books at once with a genealogical model. Filtering by county history or published genealogy, expanding a surname to its OCR variants, and combining a surname with a place in a single query are not things IA’s search supports.
- You are doing systematic research and need documentation of what was searched. Every BookTrace search produces citation-ready language recording the actual expanded terms and filters, including specific negative-search language when nothing was found.
- The question is “where in this corpus does this name appear?” rather than “does this particular book mention this name?”
- You are working to BCG or GPS standards and need to demonstrate a reasonably exhaustive search across the derivative-source literature.
The two tools compose. A typical workflow uses BookTrace to identify candidate books across the corpus, then jumps to Internet Archive to read each candidate at the estimated page. BookTrace is the card catalog; the Internet Archive is the library. Neither replaces the other, and using both is faster and more thoroughly documented than using either alone.
How the BookTrace Index Was Built
Book Selection
Books entered the index through two automated pipelines run against the Internet Archive’s public API in May and June 2026. The first pipeline queried the Archive’s genealogy collection using keyword and subject filters for genealogically relevant material. The second expanded coverage into the much larger americana collection, applying a relevance-scoring function to identify county and local histories, published genealogies, vital records abstracts, city directories, biographies, and similar material.
A book scoring above the relevance threshold entered the catalogue. A small number of harvested items (467 of 15,909) were excluded entirely, for being non-English, extremely short, or flagged during initial filtering, and never appear in any BookTrace result.
The Two Layers
BookTrace searches two layers simultaneously:
- Layer 1: Book text (extracted entities). For the 7,179 books whose OCR text could be read reliably, the pipeline processed the full text through a natural language processing model (spaCy) and extracted every text span the model classified as a person or a place. The result is an index of 39,258,145 extracted entities: 28,549,829 person spans, 10,018,095 place spans, and 690,221 other location spans, each tied to a book and an estimated page.
- Layer 2: Catalogue metadata. For all 15,442 books, whether or not their text was readable, BookTrace indexes the Internet Archive’s catalogue metadata: title, subject headings, and description.
This two-layer design is deliberate. Excluding non-OCR books entirely would hide their existence from researchers; treating them as if their text had been searched would produce a false signal. BookTrace does neither — it shows you both a book’s catalogue presence and its un-searched status.
The 53.5% Text-Searchable Gap
Of the 15,442 catalogued books, 7,179 (46.5%) have their full text in the Layer 1 index. The remaining 8,263 (53.5%) fall into three categories:
| Status | Count | What it means |
|---|---|---|
| No OCR | 6,298 | The book is a microfilm scan, a handwritten manuscript, or lacks a text layer for technical reasons. |
| OCR degraded | 1,177 | OCR was attempted, but text quality fell below a reliability threshold. Indexing degraded text produces more false extractions than valid ones. |
| Download error | 788 | The OCR file was unreachable, being access-restricted, blocked by a network failure, or missing. Some of these are retryable in a future pass. |
This gap is not a failure of the pipeline; it is a reality of the collection. Genealogical source material is disproportionately handwritten or microfilmed. What matters for your research is that BookTrace tells you, on every result card and in every coverage statement, which side of this gap a book sits on.
Using BookTrace: The Search Fields
Surnames
The surname field is the heart of the tool. Type a surname and press Enter to add it as a “chip”; you may add up to twelve. Multiple surnames support FAN-club searches (Friends, Associates, Neighbors) and family-group searches. For example, searching a research subject’s surname alongside the surnames of known associates finds books where the cluster appears together.
Spelling and OCR variant expansion is applied automatically. Each surname is expanded to include likely spelling and OCR variants (Hanks also matches Hankes, Hanx, Henks, for example), following the same principles as the VariantChronicles tool. The exact expansion used is always reported back to you in the search’s citation summary.
Wildcards are supported for researchers who want direct control: ? matches exactly one unknown letter (H?nks) and * matches any run of letters (Han*). Two rules apply. First, a term containing a wildcard is out of scope for variant expansion; the wildcard expresses your exact pattern and is searched as written. Second, terms must begin with at least one fixed letter — leading wildcards are not accepted. Note also that a term whose only wildcard is a trailing * is searched in both layers, while a term containing ? (or a * in the middle of the term) can only be searched against book text. The catalogue layer cannot be pattern-searched that way, and the results page states this directly.
Very common surnames may trigger a cap on max results. A surname like Smith can match thousands of distinct books; when a term’s matches are truncated, BookTrace notifies you directly, in both a banner and the citation summary, and suggests narrowing with places, dates, or a source type. Truncation is always disclosed.
Given Name (Optional)
An optional given-name filter narrows surname matches to books where a person entity with a matching given name was also recorded. Period abbreviations are expanded bidirectionally. Searching “William” also matches “Wm.” and “Wm”; searching “Jno.” also matches “John”. The expansion draws on a table of standard nineteenth-century written abbreviations (Thos., Chas., Geo., Jas., Robt., Benj., Margt., Eliz., and so on). This design choice works best for research in the nineteenth and early twentieth centuries, when these abbreviations were standard written practice; additional given-name matching strategies are planned for future versions of BookTrace.
Use this filter with caution, and search surname-only first. The given-name data in the index is incomplete and error-prone for structural reasons explained in the limitations section below. The filter drops books where the surname matched but the given name was not recorded or was written differently, which means it will exclude valid records. It is a narrowing tool for unmanageably large result sets, not a precision instrument. Nickname and diminutive matching (Peggy for Margaret, Polly for Mary) is not included in the current version; those mappings are genealogically consequential enough to warrant their own carefully designed release.
Places
Enter states, counties, cities, or other place names as “chips,” up to twelve. State names and abbreviations expand bidirectionally: Kentucky also matches KY and Ky, and vice versa. Place matching runs against the place entities extracted from book text (Layer 1) and against catalogue metadata (Layer 2).
Keywords in Catalogue (Optional)
This field searches the Internet Archive’s metadata about each book (title, subject, and description), never the book’s text. It is useful for terms that describe a book rather than appear in it: “county history,” “muster rolls,” “Quaker records.” Because it is a metadata-only field, it applies to all 15,442 catalogued books regardless of OCR status.
Search Scope, Proximity, and Name Logic
- Where to search. Two checkboxes, both on by default: book text (extracted entities) and catalogue entries. Unchecking one restricts the search to the other layer.
- How close together. When you search multiple terms against book text, a proximity control offers three settings: same page, within roughly ten pages, or anywhere in the book (the default). Because page numbers in this index are estimates (see below), proximity constraints are approximate by nature.
- Multiple surname matching. A toggle controls whether books must match any of your surnames (OR, the default, appropriate for FAN searches) or all of them (AND, appropriate when you specifically want books mentioning an entire cluster).
Probable Source Type
A dropdown filters results by record type: biography, city directory, county history, local history, military record, passenger list, periodical, published genealogy, or vital records abstract. Two disclosure points apply. First, the filter carries the qualifier “based on title” because classification was automated from title patterns; it is a strong heuristic but is not a librarian’s judgment. Second, roughly 6,100 books could not be auto-classified by the algorithm. When you filter by type, an inline checkbox lets you decide whether unclassified books are included or excluded, ensuring that no unclassifiable book is removed from scope without your knowledge.
Publication/Scan Year
A date-range filter is available, with an important caveat built into its label: the field is “Publication/scan year (as recorded by IA).” Some books carry dates reflecting the year the Internet Archive scanned them, not the year they were published; you will occasionally see values like 2024 or 2025 on books that are clearly a century old. Additionally, 1,544 books (10.0%) have no recorded date at all. The “include items with unknown date” toggle defaults to on. If you turn it off, BookTrace tells you how many items are being excluded, and any deliberate date filter should be acknowledged in your research notes as a scope limitation.
Interpreting BookTrace Results
The Three Result Groups
Results are presented in three fixed groups:
- Group 1: Text matches. Every found term matched in the book’s extracted text. These are the strongest results: the terms were detected inside the book itself, with approximate page numbers.
- Group 2: Split matches. Found terms were split between layers, with some matched in the book’s text and others only in its catalogue metadata. A book whose title mentions Kentucky and whose text contains Hanks would land here.
- Group 3: Catalogue matches. Every found term matched only in the catalogue metadata. For books without indexed text, this is the only kind of match possible. The term may well appear in the book’s actual text, but BookTrace cannot confirm it. Group 3 results are leads to be investigated, not text hits.
Every result set is accompanied by a coverage statement reminding you that Group 1 and Group 2 results can come only from the 7,179 text-indexed books, while Group 3 results may come from any of the 15,442 catalogued books.
The Result Card
Each result shows the book’s title, date, contributing institution, and probable source type; a colored indicator for each of your search terms showing where it was found (text, catalogue, both, or not found); approximate page numbers for text matches; and a direct link to the book on archive.org.
Two card elements deserve specific explanation:
- “Contributor” is the digitizing institution (the library that scanned the book), not the publisher, author, or subject.
- The extracted-names count on name-dense volumes. A small number of items in the corpus, such as annual Official Catholic Directory volumes, contain hundreds of thousands of extracted names and will surface constantly for common surnames. When a book’s entity count exceeds 50,000, the card states this to you (for example, “~340,000 names extracted from this volume”). BookTrace does not down-rank or hide these books — you see the reality and decide whether to review or skip. If directory-type volumes are cluttering a search, apply the source-type filter to exclude them.
What “~p. 47” Means
Every page number in BookTrace is an estimate, not a recorded page number. It is computed from each entity’s character position within the OCR text, relative to the book’s total length and page count. On typical books the estimate is accurate to roughly ±3 to 10 real pages — better for books with clean OCR and uniform page density, worse for books with unnumbered plates, dense index sections, or inserted material.
The practical rule: open the book at the estimated page, then scroll several pages in each direction. If you know the surname you are seeking, the book’s own index (when it has one) can close the gap quickly.
Sort Order
After results load, you may re-sort by date (oldest or newest first), title, or contributor. The default ordering surfaces books with stronger and more concentrated term matches first; it is labeled simply “Default order” because BookTrace does not present an internal scoring formula as though it were a meaningful, explainable ranking.
What an Extracted Entity Is, and What It Is Not
Every Layer 1 result in BookTrace rests on the same foundation: a statistical language model read the book’s OCR text and classified certain text spans as person names or place names. This is statistical classification, not verified fact, and researchers should understand its failure modes before drawing conclusions.
- False positives exist. The model occasionally labels non-names as names. Examples visible in the raw data include “surnames” like LIBRARY or Reunions, and OCR fragments stored as names. These generally produce noise you will recognize instantly, not misleading matches.
- False negatives exist. Names in inverted format (“Lake, James”), names in unusual spacing or typography, and names damaged by OCR may be missed or mis-parsed entirely.
- The surname/given-name split is positional, not linguistic. For each person entity, the pipeline split the raw text span on the last space: “James E. Lake” becomes surname Lake, given names James E. This handles most Western name formats correctly but breaks on inverted names, on suffixes (“John Lake Jr.” parses with surname Jr.), and on titles (“Mrs. John Lake” parses with given names Mrs. John). This is the structural reason the given-name filter is best-effort and surname-first searching is the recommended workflow.
Every Layer 1 result should therefore be read as: the string was detected as a name or place by an automated model in this book’s OCR text, at approximately this page. It is a signal to open the book and read the actual passage.
No Snippets, and No Sentence-Level Co-occurrence
Two further limitations are important enough to state separately.
BookTrace cannot show you the passage. The index stores each extracted name and its approximate page, not the sentence it appeared in. Unlike most search interfaces, BookTrace cannot display a text snippet around your match. To read the passage, follow the link to the Internet Archive and navigate to the approximate page. Snippet display is a named future enhancement.
Same-page co-occurrence is not sentence-level co-occurrence. When BookTrace reports that a surname and a place both appear on approximately the same page, it means both were independently extracted from that region of the text, within a few dozen lines of one another, no more. It does not mean they appeared in the same sentence, or that the place describes the person. Treating a same-page match as evidence of association without reading the passage is an evidentiary error the tool’s design cannot prevent; only the researcher can.
Negative Searches and GPS Documentation
For researchers working to the Genealogical Proof Standard, a well-documented negative search is as valuable as a positive one. BookTrace distinguishes three zero-result states, and the distinction matters:
- True negative. Zero matches in book text and zero matches in catalogue metadata. This is the strongest negative signal BookTrace can produce: across all 15,442 books, neither the indexed text of the 7,179 text-searchable books nor the metadata of any catalogued book contained your terms. BookTrace generates negative-search citation language for this state, suitable for direct paste into your research log.
- Partial negative, not text-searchable. Zero text matches, but catalogue matches exist among books whose text is not indexed. Your terms may appear in those books’ actual pages; BookTrace cannot confirm or deny it. These books are your follow-up list. Open them on archive.org and search or browse them individually.
- Partial negative, narrow scope. Your filters (date range, source type, layer selection) may have excluded books that would otherwise have matched. Any deliberate filter is a source of hidden negatives, and BCG-standard documentation should record the filters used. The citation summary generated with every search records them for you.
One boundary must be stated plainly: even a true negative in BookTrace is not a negative for the Internet Archive as a whole. BookTrace catalogues only the genealogy- and americana-relevant books that scored above a relevance threshold during index construction. It is not a search of the Archive’s 40+ million items, and a research log entry should describe it accordingly.
Citing a BookTrace Search
Every search, positive or negative, produces a citation-ready summary stating exactly what was searched: the terms as entered, the variant expansions actually applied, the filters in effect, and the scope of the index. A typical citation in BCG/GPS style:
BookTrace (evidencetoolbox.com/tools/booktrace), search for surname “Hanks” and place “Kentucky,” accessed [DATE]; searched both book text and catalogue metadata across 15,442 catalogued Internet Archive books, of which 7,179 have full text indexed.
Because the summary records the actual expanded terms, your research log captures not just what you intended to search but what was in fact searched.
A Recommended Workflow
- Search surname-only first, with default settings. Review the size and shape of the result set before narrowing.
- Narrow with places before given names. Place filtering is structurally more reliable than given-name filtering in this index.
- Add the given-name filter only when a result set is unmanageably large, and treat the books it removes as unexamined, not eliminated.
- Work Group 1 and Group 2 results first. These have page pointers. Open each book at the estimated page and read outward.
- Treat Group 3 results as a follow-up list. Open each on archive.org and use the Archive’s own in-book search or the book’s index, since BookTrace has not read these books’ text.
- For elusive OCR-damaged names, try wildcards (
H?nks,Han*) after the variant-expanded search, remembering that wildcard terms bypass variant expansion and, except for trailing-*patterns, search book text only. - Copy the citation summary into your research log for every search, including, especially, the negative ones.
- Verify everything in the source. Every BookTrace result is a pointer to a book on the Internet Archive. The evidence is in the book, not in the index.
Data Handling and Relationship to the Internet Archive
BookTrace stores no book content. Its index contains extracted name and place strings, approximate page positions, and catalogue metadata; every result links directly to the original item on archive.org, where the page images are served by the Internet Archive itself. BookTrace is an independent research tool built by Glenside Digital LLC and is not affiliated with, or endorsed by, the Internet Archive.
For a fuller technical account of how the index was constructed, the OCR-quality thresholds, and the complete statement of known limitations, see the BookTrace transparency report.
Index figures in this guide (15,442 books catalogued; 7,179 full-text indexed; 39,258,145 extracted entities) reflect the index as of July 2026 and will be updated as the corpus expands.
About the Author
Nathan is an avocational genealogist and the founder of Evidence Toolbox. His research practice is grounded in the Genealogical Proof Standard, with primary-source work conducted at major repositories including the Library of Congress. He builds the tools on this site to solve problems encountered in his own research, and field-tests each one against real family lines before release. You can reach him at contact@evidencetoolbox.com.