Hi everyone,
I am the original author of Searx. I started Hister with a similar motivation: reducing our dependence on external search engines while keeping searches and personal data under our control.
Searx is a metasearch engine that forwards queries to other search providers. Hister takes a different approach. It builds a private full text index from content you choose, then searches that index entirely on your own infrastructure.
Hister can automatically index pages through its Firefox and Chrome extensions. It can also watch local directories, import browser history and bookmarks, index individual URLs, and crawl complete documentation sites.
The feature I find most useful is offline previews. Hister stores the readable content and HTML of indexed pages locally. You can open a result in a clean and sanitized preview beside the search results without visiting the original website again.
Some other features:
- Full text search across web pages, PDFs, docx files, Markdown, OrgMode and text files
- Phrase searches, field filters, date filters, wildcards, negation, aliases, labels, facets, and result priorities
- Optional semantic search using an embeddings endpoint you configure
- Persistent website crawls
- Imports from browser history, Linkwarden, Karakeep, Shaarli, Wallabag, and Linkding
- Web, terminal, command line, HTTP API, and MCP interfaces
- SQLite and PostgreSQL support, plus optional multiple user hosting
Hister cannot replace a global search engine (yet) for subjects you have never encountered because it only searches what you have indexed. My workflow is to search Hister first, then use its shortcut to fall back to traditional search when I need broader web results.
The project is free software under the AGPLv3+ license. It can be installed as a standalone binary or with Docker.
Project: https://github.com/asciimoo/hister
Website and documentation: https://hister.org/
Small read-only demo: https://demo.hister.org/
I'd appreciate feedback, questions, and suggestions as well as joining our growing community.
AI disclosure: AI assisted contributions are not strictly prohibited, but all contributions should be made by humans. More details: https://github.com/asciimoo/hister/blob/master/CONTRIBUTING.md#ai-policy
The summary has numerous inaccuracies. Most importantly: it is pretty easy to delete content by topic or age. The
hister deletecommand can accept a search query to remove only matched documents. The same is true on the web UI "actions -> remove all matching documents". You can quickly filter by age, simply queryupdated:>365d. Combine it with URLs, labels, domains or phrases.As I wrote, you can simply automate deletion by document age. Schedule a delete event on each day with the desired retention time defined as a filter expression. Database size limit isn't available yet.
No, sorry, I don't have time to correct a copy of a multiple screens long AI prompt.
Which parts are confusing?
I think it is more abrasive to copy/paste a poorly formatted LLM output instead of taking the time and summarizing it to a few sentences just as you did in your previous post.
Feedback or questions is appropriate. We don't need to come in abrasive, and you didn't need chatgpt to ask a followup, so we don't need to be antagonistic.
This is one of the problems with relying on AI... it can produce an overwhelming amount of content with errors and inaccuracies throughout. If you don't review and know the content yourself, you won't know what it got wrong.
It's genuinely rude to lazily use AI to produce such a large babble of details and then ask someone else to review it for errors when you haven't reviewed it yourself. I know you didn't mean it that way but that's nonetheless the result. People are going to read your comment and be misled about this project all because they assume AI is accurate and you didn't review its results.
Edit: Sheesh... I hadn't even got to your shitty comments that followed. You use AI to make a low effort but highly verbose post and then get mad that the repo author won't review it in detail for you when you can't be arsed with reading the docs yourself.
Yep, and this induces anyone who wants to review the resulting slop to need to turn to AI as well to even deal with the amount (or dismiss as a whole).
AI is a viral infection on software development.
Yep... couldn't agree more. I has its uses but too many people are using it and disconnecting their brain. Going to meetings where people have used it to determine requirements, do analysis or even transform data and then don't even review the results themselves before presenting and asking us to review them is rage-inducing. "What do you mean it has problems? Where? What's wrong??" And there's like 5 things I've spotted in 2 minutes and it's clear they've not even reviewed it themselves.
This guy and his long-ass, poorly formatted AI-generated "documentation" no one asked for and then after numerous errors are pointed out.... "Will you review the rest?" The audacity!! 😄