Tally

Methodology and data sources

Fund fee data

Fee figures come from the SEC's Mutual Fund Prospectus Risk/Return Summary data sets: structured XBRL facts that every open-end fund and ETF must file with its prospectus. We read the net expense ratio (after waivers), gross expense ratio, 12b-1 distribution fee, maximum front-end and deferred sales charges, and portfolio turnover for each share class, keeping the most recently filed value. Share classes are mapped to tickers with the SEC's company_tickers_mf.json and named with the Investment Company Series and Class dataset.

The database currently holds 26,569 funds, 26,546 of them from SEC filings (latest data date 2026-08-31), plus a curated seed list of widely held funds with their tracked index and category. Funds not found in either are flagged, and you can enter the expense ratio from the prospectus by hand.

Refresh the data with npm run ingest:sec. The SEC requires a descriptive User-Agent; set SEC_USER_AGENT.

Categories and alternatives

Every fund is placed into one of about sixty categories (US large growth, emerging markets, core bond, 2050 target date, and so on). Seed funds are categorized by hand; SEC-sourced funds by rules on the fund name; anything unrecognised can be classified by the AI model from its name. Each category lists broad, low-cost index funds that deliver the same market exposure. A fund is matched to those champions, preferring one that tracks the identical index, then the same category; only alternatives cheaper than the current fund are shown. Target-date funds are matched to index target-date series with the same target year.

This is a category match, not a holdings or return-correlation match. An active fund and its index alternative will not move identically. The premise, well supported by SPIVA and Morningstar's Active/Passive Barometer, is that cost is the most reliable predictor of a fund's future relative performance within a category.

Cost and projection math

Annual fund cost is market value times net expense ratio. The advisory fee, if any, is applied to the whole account. Projections run year by year: contributions are added at the start of the year (less any front-end load, if enabled), the balance grows at the chosen gross return, and fees are deducted on the average balance. The "no fees" line is the same path with zero cost and is a theoretical ceiling, not an achievable outcome.

Not modeled: taxes triggered by switching in a taxable account, bid-ask spreads, cash drag, securities-lending income, differences in the funds' actual returns, and sales loads already paid (those are sunk). Returns are nominal and constant; real markets are not.

Statement parsing

PDF text is extracted inside your browser with pdf.js; the file itself is never uploaded. Names, account numbers, SSNs, tax IDs, dates of birth, emails, phone numbers and addresses are replaced with placeholders, then you review and can edit the exact text, mask more terms, or opt out of AI reading entirely. If you send it on and an Anthropic API key is configured, that text goes once to Claude with a strict output schema to extract positions, values and any advisory fees; the model is told the text is data, never instructions, and its free-text output is length-limited. Otherwise the text is scanned on the server for known fund tickers and nearby dollar amounts, which is less reliable. Scanned image PDFs have no text layer and cannot be read this way; type the holdings in instead. Fund names that are not in the SEC database may be sent to the model to be categorized, only if AI was allowed on the review or manual-entry step; shared links never trigger AI calls. Tally stores none of this text; Anthropic processes it under its own data policy. Requests are rate limited per address and nothing is written to disk or a database.

Shared links

A share link shows the same analysis rescaled to a sample $100,000 portfolio. Real balances, the statement date, account labels and share counts are removed before the link is created; the funds, their weights, each fund's fees and the advisor fee are kept, so every percentage and every ratio is exact while the dollar figures are simulated. The link carries that redacted data itself after the # in the address, so browsers never send it to the server when opening the page; it is signed with an expiry seven days out. Nothing is stored on our servers, and once the week is up the link stops working. Shared pages ask search engines not to index them, and because anyone with the secret could mint a link, viewers are told that Tally did not verify who created it.

Fee benchmark (opt-in)

The benchmark is built only from results people chose to add, one unticked checkbox at a time, after seeing the exact record. A record holds the advisor firm, the advisory fee rounded to 0.05 percentage points, a balance range (under $50k, $50k to $250k, $250k to $1M, over $1M), an optional US state, and the funds by ticker with their portfolio weights and expense ratios, plus the date to the day. It never holds names, account labels, dollar amounts, hand-typed fund names, or your address; the server does not log the sending address either, and rate limiting uses a hash that changes daily. Groups with fewer than 10 records are not published, in any cut. Each submission returns a random receipt id, shown once, which removes the record for 30 days. Aggregates are medians and quartiles recomputed at most hourly. The sample is self-selected and should be read as "what Tally users who opted in pay", not as a survey of all investors.

Not advice

Tally is an educational tool. It does not know your tax situation, goals or risk tolerance, and it does not recommend buying or selling anything. Use it to ask better questions.