Skip to content

Fix: content scraping and filtering bugs (noscript, preserve, pre/code) - #2296

Open
dajiaohuang wants to merge 1 commit into
unclecode:developfrom
dajiaohuang:bugfix/content-scraping-filters
Open

dajiaohuang wants to merge 1 commit into
unclecode:developfrom
dajiaohuang:bugfix/content-scraping-filters

Conversation

@dajiaohuang

Copy link
Copy Markdown

Summary

Fixes three related content scraping and filtering bugs:

Changes

File Change
crawl4ai/content_scraping_strategy.py Add _balance_noscript_tags() to balance malformed noscript tags before lxml parsing
crawl4ai/content_filter_strategy.py Skip preserved tags in _remove_unwanted_tags(); skip pruning in pre/code blocks
crawl4ai/content_filter_strategy_lxml.py Accept preserve params; skip preserved tags during strip/metrics/pruning; never prune pre/code children

Validation

- unclecode#2293: Balance unclosed <noscript> tags before parsing to prevent body content loss when nested/malformed noscript is present
- unclecode#2125: Respect preserve_tags/preserve_classes when removing excluded tags in both PruningContentFilter and PruningContentFilterLXML
- unclecode#2110: Skip pruning inside <pre>/<code> blocks to preserve whitespace-only spans in syntax-highlighted code, preventing corruption in fit_markdown
- Add preserve_classes/preserve_tags parameters to PruningContentFilterLXML for parity with the BeautifulSoup version
Sign up for free to join this conversation on GitHub. Already have an account? Sign in to comment

Labels

None yet

Projects

None yet

Development

Successfully merging this pull request may close these issues.

1 participant