↓Skip to main content
  1. Blog/

Revisiting Hister

·4 mins

In posts about SearXNG on Hacker News, the original creator of SearX showed up in each post to promote his new project, Hister. Hister is essentially a personal search engine, based on the premise that instead of blindly indexing and crawling the web, most of the content we need to reference is content we’ve already seen before, and we should index and reference that as a first-class set of primary sources. The primary way that the index is built in Hister is through browser extensions that quietly archive the pages that one visits. [1] MCP is a big [new I guess?] selling point as it’s essentially meant to also expose the content to agents.

For the first couple of months after setting it up, I mostly used it to find Hacker News and Kagi Small Web posts whose titles I’d forgotten. This worked as long as I could remember some words from the post. I didn’t have much reason to open Hister otherwise.

I should have realized earlier that I could import collections into it. For example, they offer a few collections you can import like the Go Language reference and all the RFCs.

For some reason, I started thinking about a much more complicated project originally: scraping a corpus of award-travel writing and queueing batch processing from LLMs to extract and build a database of selected data points I might otherwise forget. The problem I was trying to solve I guess is best illustrated by an example: someone pointed out to me that British Airways flights booked with Cathay Asia Miles have much lower taxes and fees, something travel blogs covered adequately since 2019 but that had become a blind spot for me, because well, I hadn’t thought about it for years since that first became public knowledge. I originally wanted to be able to query for interesting knowledge about British Airways so I could stop myself from making a mistake or forgetting a creative approach that would cost me time or a couple hundred dollars. Essentially improve the signal to noise ratio for searching up info about award travel since most of these articles contain little besides travel-blog marketing copy about one of four premium US travel credit cards.

At some point while away from my computer, I realized that there was no reason for me to pre-process potentially 100,000+ blog posts. I just needed to import these articles into Hister as a collection and let an agent essentially just pull from it when necessary. The cost of searching and processing this data would’ve been considerably lower than doing huge batch jobs with a constant need for finetuning.

Hister does actually have a built-in web scraper, but it seemed quite basic (it recursively indexed from a starting URL). I knew most of these blogs had a fairly complete sitemap.xml exposed so I ended up rolling a persistent scraper focusing on archiving directly from sitemap.xml.[2]

Hidden in the documentation (at least to me) was that it was possible to enable semantic search with an embedding model. I currently have a synthetic.new subscription, which coincidentally offers unlimited embedding with the hf:nomic-ai/nomic-embed-text-v1.5 model (embedding is cheap elsewhere though, I can always migrate).

I finally connected the Hister MCP to my OpenCode, and so far it’s been working well when I ask about what my Hister Index says about a few niche award travel problems I’ve come across (e.g. booking with obscure airlines and the availibility of certain routes). It isn’t a perfect solution (not sure I’ll ever find one for this vague problem shape) but seems a lot better than my original approach at least.

I’ve indexed about 140,000 pages so far. I did initially have some growing pains with memory. The default Proxmox community script gave Hister about 1.5GB of RAM and 512MB of swap, so it started crashing when I searched. Upping that to 4 GB of RAM seems enough; I added another 4 GB of swap too. This is all with SQLite too, so far I haven’t needed to migrate to Postgres.

Anyways, I think discovering that I could’ve done all of this without building a custom pipeline and writing a ton of additional code has been a relief. I’m also glad I found a genuine use case that lets me properly justify keeping Hister around, rather than treating it as a fun toy and yet another exercise in data hoarding.

[1]: I set various Hister rules to avoid indexing content that related to myself and my partner’s personal dealings (i.e. containing my name).

[2]: Just to be clear, I did not hammer these sites and set a reasonable crawl delay on each of them. A few of these sites also did not allow any crawling at all unfortunately, and I wasn’t interested in investing time circumventing Cloudflare so I moved on without them.