Fix: content scraping and filtering bugs (noscript, preserve, pre/code) - #2296
Open
dajiaohuang wants to merge 1 commit into
Open
dajiaohuang wants to merge 1 commit into
dajiaohuang wants to merge 1 commit into
Conversation
- unclecode#2293: Balance unclosed <noscript> tags before parsing to prevent body content loss when nested/malformed noscript is present - unclecode#2125: Respect preserve_tags/preserve_classes when removing excluded tags in both PruningContentFilter and PruningContentFilterLXML - unclecode#2110: Skip pruning inside <pre>/<code> blocks to preserve whitespace-only spans in syntax-highlighted code, preventing corruption in fit_markdown - Add preserve_classes/preserve_tags parameters to PruningContentFilterLXML for parity with the BeautifulSoup version
This file contains hidden or bidirectional Unicode text that may be interpreted or compiled differently than what appears below. To review, open the file in an editor that reveals hidden Unicode characters.
Learn more about bidirectional Unicode characters
Sign up for free
to join this conversation on GitHub.
Already have an account?
Sign in to comment
Add this suggestion to a batch that can be applied as a single commit.This suggestion is invalid because no changes were made to the code.Suggestions cannot be applied while the pull request is closed.Suggestions cannot be applied while viewing a subset of changes.Only one suggestion per line can be applied in a batch.Add this suggestion to a batch that can be applied as a single commit.Applying suggestions on deleted lines is not supported.You must change the existing code in this line in order to create a valid suggestion.Outdated suggestions cannot be applied.This suggestion has been applied or marked resolved.Suggestions cannot be applied from pending reviews.Suggestions cannot be applied on multi-line comments.Suggestions cannot be applied while the pull request is queued to merge.Suggestion cannot be applied right now. Please check back later.
Summary
Fixes three related content scraping and filtering bugs:
[Bug]: cleaned_html drops the page body when it is nested inside an unclosed <noscript> #2293: drops page body when nested/unclosed
<noscript>tags are present (common with LiteSpeed Cache + Google Tag Manager lazy-load). Added_balance_noscript_tags()pre-processor that inserts missing closing tags before parsing to match browser scripting-enabled behavior.[Bug]: preserve_tags/preserve_classes are a no-op for excluded tags (aside, nav, footer, header, form) #2125:
preserve_tags/preserve_classesoptions were silently ignored for excluded tags (aside, nav, footer, header, form). Fixed_remove_unwanted_tags()to skip preserved elements in both BeautifulSoup and LXML implementations. Also added the missingpreserve_classes/preserve_tagsparameters toPruningContentFilterLXMLfor API parity.PruningContentFilter drops whitespace-only spans in <pre>/<code> — #1181's bug, still present on the fit_markdown path #2110:
PruningContentFilterremoves whitespace-only spans inside<pre>/<code>blocks (used by syntax highlighters), corrupting code formatting infit_markdown. Added guard to skip scoring/pruning entirely for pre/code subtrees, matching the fix already applied toremove_empty_elements_fast.Changes
crawl4ai/content_scraping_strategy.py_balance_noscript_tags()to balance malformed noscript tags before lxml parsingcrawl4ai/content_filter_strategy.py_remove_unwanted_tags(); skip pruning in pre/code blockscrawl4ai/content_filter_strategy_lxml.pyValidation