| Age | Commit message (Collapse) | Author |
|
|
|
|
|
This has only been needed because we were flagged by HTTP2 fingerprinting.
Now, since we use curl cffi, we can bypass the fingerpinting, so we can
use HTTP2 just fine without getting blocked.
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
requires https://github.com/searxng/searxng/pull/6620
|
|
Signed-off-by: Markus Heiser <markus.heiser@darmarit.de>
|
|
|
|
|
|
|
|
|
|
|
|
Closes: https://github.com/searxng/searxng/issues/6545
|
|
|
|
|
|
- Closes: https://github.com/searxng/searxng/issues/6622
Signed-off-by: Markus Heiser <markus.heiser@darmarit.de>
|
|
|
|
The previous implementation ran into an error if the search term contained words
other than just the location (ValueError was raised).
To test engine use search terms like:
!ddw weather berlin germany
Signed-off-by: Markus Heiser <markus.heiser@darmarit.de>
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
alphabet
|
|
random.choices
|
|
|
|
Removing the presearch engine and configuration because it got shutdown.
ref.: https://news.presearch.io/a-message-from-the-presearch-team-aaa448052b2b
https://x.com/presearchnews/status/2080744472791441807
|
|
|
|
|
|
dict
|
|
|
|
The code used here has always been "bad" because `about` shouldn't be used
as data source, but the engine probably broke when type checks / dataclasses
for the about parameter in engines has been added with
<https://github.com/searxng/searxng/pull/6258>.
Error log:
```
WARNING searx.engines.public domain im: ErrorContext('searx/engines/public_domain_image_archive.py', 143, '\'url\': _clean_url(f"{about[\'website\']}/images/{result[\'objectID\']}"),', 'TypeError', None, ("'EngineAbout' object is not subscriptable",)) False
ERROR searx.engines.public domain im: exception : 'EngineAbout' object is not subscriptable
Traceback (most recent call last):
File "/home/bnyro/Projects/searxng/searx/search/processors/online.py", line 253, in search
search_results = self._search_basic(query, params)
File "/home/bnyro/Projects/searxng/searx/search/processors/online.py", line 239, in _search_basic
return self.engine.response(response)
~~~~~~~~~~~~~~~~~~~~^^^^^^^^^^
File "/home/bnyro/Projects/searxng/searx/engines/public_domain_image_archive.py", line 143, in response
'url': _clean_url(f"{about['website']}/images/{result['objectID']}"),
~~~~~^^^^^^^^^^^
TypeError: 'EngineAbout' object is not subscriptable
```
|
|
Qwant now requires a `datadome` cookie that it returns
in the first search response as a `Set-Cookie`.
This cookie has to be sent for all requests, otherwise they
will be blocked.
This means that now, the first search request is blocked (results in CAPTCHA),
and only the subsequent searches work (same happens on the Qwant website for me).
However, I don't think it's worth repeating the same
search request multiple times very quickly because
that also makes us more suspicious.
|
|
Changes:
- the `embed.js` request now requires a user agent header
- we include a user agent in the "actual" request (I dropped it by accident)
- we only send the first 4 decimal places of the location
instead of 7+ (not required, but harder to detect)
|
|
|
|
where possible (#6394)
Refactor engines that parse ISO 8601 dates with strptime to use
fromisoformat instead. In most cases this is a direct replacement of
strptime(text, "format") with fromisoformat(text).
For engines where the source has a trailing "Z" that strptime consumed
as a literal (e.g. "%Y-%m-%dT%H:%M:%S.%fZ" in huggingface.py), add
rstrip("Z") to keep the output naive and preserve the existing behavior.
In sogou.py the date is extracted with a regular expression, which can
yield strings like "2026-7-11". strptime accepts this via its format
string, but fromisoformat does not. To preserve the existing behavior
and satisfy the format fromisoformat expects, add zero-padding for the
month and day.
Closes: #6098
---------
Signed-off-by: OneVth <onebrotravel@gmail.com>
|
|
|
|
Heexy now passes the cacheft token via cookies and no
longer via HTTP headers.
Hence, the engine is broken without that change.
|
|
|
|
I was testing the new Kagi engine and found for some queries I was getting a
`KeyError` exception from the result parsing.
This PR ensures we only assumes the `url` key exists, and we use `.get()` to
retrieve values for keys that may not be present.
Their API documentation [1] clarifies that only the `url` and `title` properties
are required/guaranteed in the search result object, but this is not correct:
File "/home/patrick/code/searxng/searx/engines/kagi.py", line 161, in response
title=html.unescape(result["title"]),
~~~~~~^^^^^^^^^
KeyError: 'title
Heard back from Kagi support:
> The image results "title' should be marked as optional, as many images likely
> don't have titles - as you've noticed. The only required field there should be
> the "url" field.
[1] https://kagi.redocly.app/api/docs/openapi/search/search#search/search/t=response&c=200&path=data/search
[2] https://github.com/searxng/searxng/issues/2247#issuecomment-4692976877
|
|
|