mirror of
https://github.com/crewAIInc/crewAI.git
synced 2026-08-10 16:32:28 +00:00
Five real defects from Bugbot, none of them cosmetic. Tool-scoped policy never applied (high). `resolve_tool_failure_policy` read `tool_failure_policy` off the object handed to it, but every execution path passes the `CrewStructuredTool` wrapper, which never carried the attribute -- and `BaseTool` never declared it in the first place. A tool-scoped `raise`/`ignore` was silently ignored while the docs and a unit test claimed otherwise; the test passed only because it called the resolver directly with an authored tool. Declared the field on `BaseTool`, propagated it through `to_structured_tool()` and `CrewStructuredTool`, and made resolution fall back through `_original_tool` so either shape works. A failed call still printed the green "Completed" panel, then the red one. That is the terminal version of the exact bug this PR is about. Suppressed the success panel when the call reported failure. A raised tool printed twice: `ToolUsageErrorEvent` already renders a red panel, and the new failure panel repeated it. The event is still emitted -- policy and traces need it -- but the duplicate console output is gone. Both decisions now live in named predicates on `ConsoleFormatter` rather than inline in the listener closure, so they are directly testable. Unknown tools were reported on the ReAct path but silently ignored on all three native paths, so the same miss was loud or silent depending on executor style. Native paths now record `UNKNOWN_TOOL` too. This also surfaced a live `NameError`: ruff had pruned `ToolFailureReason` from `agent_utils` as unused, so the new branch would have crashed at runtime. `LiteAgentOutput` had `tool_failures` but not `has_tool_failures`, which the PR promised on all three output types -- an `AttributeError` for any caller sharing one check across result types. Testing: 16 further tests, 45 total. Two console tests were passing vacuously because `emit()` dispatches sync handlers on a thread pool, so the assertions raced the handler; they now assert on the predicates directly, and the native-path test drains the bus with `flush()` and checks the synchronously-written record. Full suite still matches baseline exactly at 377 pre-existing failures. Co-Authored-By: Claude Opus 5 (1M context) <noreply@anthropic.com> Claude-Session: https://claude.ai/code/session_01ETacm2dMASfpMAYUiDu5YG