mirror of
https://github.com/crewAIInc/crewAI.git
synced 2026-09-20 01:55:38 +00:00
* fix(tools): read octet-stream URLs by sniffing the body URLReadTool resolved content type from the Content-Type header and then the URL path extension. Presigned object-store links carry neither: they pin every object to application/octet-stream and use a content hash for a path, so a SharePoint download landing in R2 was refused outright. Sniff the already-fetched body as a third source, consulted only after the header and both URL extensions come back with nothing. The sniff can turn a refusal into a read but never a read into a different read, so no URL that works today changes behavior. Fails closed: a zip is DOCX only when word/document.xml is in its central directory, so an .xlsx keeps its honest refusal instead of surfacing a misleading "failed to read DOCX"; text requires a strict, whole-body UTF-8 decode with no NUL byte; an empty body identifies nothing. Co-Authored-By: Claude Opus 5 (1M context) <noreply@anthropic.com> * feat(tools): extract text from XLSX URLs The reported presigned SharePoint link is a spreadsheet, so sniffing the body identified it as OOXML but still had nowhere to send it: URLReadTool had no XLSX extractor, and the file would have been refused even with a correct spreadsheetml Content-Type. Read workbooks with openpyxl, already a core crewai dependency, so this adds no new one. Sheets are emitted as CSV under a "Sheet <name>:" heading, mirroring the PDF extractor's per-page shape. read_only streams the sheets instead of building the whole object graph and data_only takes cached values, both of which matter for a workbook arriving from an untrusted URL. Cells are written through csv rather than joined, so a comma, quote or newline inside a cell cannot corrupt the grid, and trailing phantom rows are trimmed because Excel reports sheet dimensions generously. Co-Authored-By: Claude Opus 5 (1M context) <noreply@anthropic.com> * fix(tools): bound xlsx expansion and refuse ambiguous ooxml packages Bot review found two real defects in the XLSX extractor, both reproduced. openpyxl pads every row up to a sheet's declared dimension, so a single stray cell far down the sheet turned a 4.8 KB upload into 100,000 rows and 200,000 cells. Trimming only trailing blanks did not help, because the stray cell sits at the end and keeps the last row non-empty. Blank rows are now skipped as they stream, and a cell budget caps what any one workbook can hand an agent -- announced in the output rather than silently applied. A zip carrying both word/document.xml and xl/workbook.xml was classified as DOCX. Two identities is not a positive identification, so it is refused. Co-Authored-By: Claude Opus 5 (1M context) <noreply@anthropic.com> * fix(tools): keep whitespace-only xlsx cell values Bot review, verified: openpyxl's row padding arrives as None, so testing cells for exactly-empty drops it just as well as .strip() did while leaving a row whose cells the author really did fill with spaces. And rstrip() on the rendered grid removed a trailing space from the final cell along with the line terminator; only the terminator should go. Co-Authored-By: Claude Opus 5 (1M context) <noreply@anthropic.com> * fix(tools): bound xlsx scan work, not just emitted cells The cell budget only counted cells that reached the output, and blank rows skip before that point. A sheet can declare Excel's maximum dimension while holding two real cells; openpyxl then pads every row out to 16,384 columns and yields one row per gap. Measured: a 4,848-byte workbook drove 1.64 billion cell normalizations in 15.2 seconds with the budget never touched. Charge a separate scan budget per row, before the row is normalized and before the blank check, so the work a hostile sheet can demand is bounded whether or not any of it is emitted. The regression test asserts the read completes in under 5 seconds and is mutation-verified: dropping the per-row charge takes it back to 26 seconds. Co-Authored-By: Claude Opus 5 (1M context) <noreply@anthropic.com> * fix(deps): clear the six pip-audit advisories gitpython 3.1.58 has PYSEC-2026-3785 through -3788, fixed in 3.1.59; the lock now takes 3.1.61. Its exclude-newer-package cutoff is dropped rather than bumped -- the global 3-day cutoff has long since passed 2026-08-05, so that per-package pin was only holding the fix back. snowflake-sqlalchemy 1.10.0 has GHSA-8g6f-qw9x-4q6q (SQL injection and local file disclosure), fixed in 1.11.0. unstructured 0.18.32 has GHSA-4mvj-m6j5-pmf7, a full-read SSRF via the url= argument of partition(). The patched 0.24.0 requires Python >=3.11 while crewai-tools supports 3.10, so the floor carries a marker and 3.10 stays on the old line. 0.24+ also requires beautifulsoup4>=4.14.3, so the bs4 pin widens from ~=4.13.4 to >=4.13.4,<5 -- a widening, so no existing install breaks. uv resolves bs4 4.13.5 on 3.10 and 4.15.0 on 3.11+. pip-audit locally: "No known vulnerabilities found, 5 ignored", with no new --ignore-vuln entries. Only crewai-tools[xml] grows, gaining spacy and openai-whisper transitively through unstructured's extras. Co-Authored-By: Claude Opus 5 (1M context) <noreply@anthropic.com> * fix(tools): narrow bs4 find_all results without a cast Widening the beautifulsoup4 pin let uv resolve 4.15.0 on Python 3.11+ while 3.10 stays on 4.13.5, because the old unstructured line holds it back there. 4.15 types find_all precisely, so cast(Tag, link) became redundant and mypy failed the 3.11-3.13 type-checker jobs while 3.10 passed. isinstance narrowing is correct under both versions and is what AGENTS.md asks for anyway. Verified by running mypy against 4.15.0 and again against 4.13.5: browser_toolkit is clean under both, leaving only the pre-existing errors in crewai/rag/embeddings/providers/ibm. Co-Authored-By: Claude Opus 5 (1M context) <noreply@anthropic.com> * fix(deps): declare security floors in crewai-tools, not only as overrides Bot review caught a regression I introduced. override-dependencies replace the whole requirement including its marker, so gating the unstructured override on python_version >= '3.11' dropped the dependency outright on 3.10: the lock held only 0.24.1, never the 0.18 line the comment claimed. crewai-tools[xml] would have installed no unstructured at all there. Move the floors into lib/crewai-tools/pyproject.toml, where a marker split means what it says -- >=0.24.0 on 3.11+, >=0.17.2 below -- and drop the root override for unstructured entirely. The lock now carries both 0.18.32 and 0.24.1 under complementary markers. Same reasoning applies to the other two, per the nltk precedent already in that file: a uv override only shapes this workspace's lock, so consumers installing crewai-tools[snowflake] or [github] were still getting the vulnerable floors. Declared there now as well. Also documents the tool as a fit for presigned and share links from S3, R2, Google Drive, OneDrive and SharePoint -- the case this PR fixes -- while saying plainly that it reads a URL and does not authenticate. Co-Authored-By: Claude Opus 5 (1M context) <noreply@anthropic.com> * chore: update tool specifications --------- Co-authored-by: Claude Opus 5 (1M context) <noreply@anthropic.com> Co-authored-by: github-actions[bot] <41898282+github-actions[bot]@users.noreply.github.com>