if you looking for a good alternative for google search, try "SearXNG"

ColdWater@lemmy.ca · edit-2 2 days ago

if you looking for a good alternative for google search, try "SearXNG"

Mojeek Search Engine@lemmy.ml · 1 day ago

Reddit doesn’t allow us to crawl: https://www.reddit.com/robots.txt

notfromhere@lemmy.ml · 15 hours ago

Is that legally binding? What happens of they catch you, ban your IPs then you’re in the same situation as now. Literally no reason to not do it IMO.

Mojeek Search Engine@lemmy.ml · 15 hours ago

IP already hits a wall, also better to not get a reputation as a bad bot, it’s taken a while to get known for being friendly and respecting rules, to us you should follow robots

notfromhere@lemmy.ml · 15 hours ago

I seem to recall creative ways to index things without robots, e.g. browser addon that users opt into to send pages and such, essentially crowdsourcing the indexing. Anyways good to see you’re taking the high road!

Mojeek Search Engine@lemmy.ml · 8 hours ago

our preference is always to find out why the block is happening and try to convince people it should be otherwise; widespread abuse of robots.txt does no-one any good, having been crawling and indexing for so long it’s a standard that we understand and are quite fond of

we can see some of the perils and pitfalls of it too, but web builders need to be given some tools and assurances that those tools will work for them

notfromhere@lemmy.ml · 8 hours ago

That makes sense. One thing I’ve noticed with Mojeek search results compared to Google is that I do not encounter the “old web” any more on Mojeek than on Google. Are you not crawling/indexing web 1.0 blogs and sites at all?