mirror of
https://github.com/crewAIInc/crewAI.git
synced 2026-08-10 16:32:28 +00:00
CI caught a type error I should have: widening `agent` to accept a `LiteAgent` (so a standalone LiteAgent resolves its own policy) left the declared signatures behind. Widened `execute_tool_and_check_finality`, its async twin, and `ToolCallHookContext` to `Agent | BaseAgent | LiteAgent | None`, which is what those actually receive now. Seven CodeRabbit findings, all verified against the code first: `raise` was being downgraded by three enclosing handlers. With `max_execution_time` set, `_execute_with_timeout` wrapped every exception in `RuntimeError`, so `_check_execution_error` no longer recognized the passthrough and sent the task through the retry loop instead of aborting. `StepExecutor.execute` turned it into `StepResult(success=False)` and let the plan continue. `LiteAgent.kickoff` ran it through `handle_unknown_error` and printed "This is likely a bug - please report it" for what is a deliberate, configured stop. Failure records were dropped on two paths. `reset_tool_failures()` only ran in `_prepare_task_execution`, so `Agent.kickoff()` / `kickoff_async()` — which enter through `_prepare_kickoff` — accumulated records across runs. And a guardrail retry calls `execute_task` again, which resets the agent, so a tool that failed on a blocked attempt vanished from the final output entirely: a run could report zero failures having demonstrably failed one. Failures now accumulate across guardrail attempts. Writing the tests for that surfaced a further miss of my own: `Agent.kickoff()` builds its `LiteAgentOutput` in `agent/core.py` via `AgentExecutor`, not through `LiteAgent`, so `tool_failures` was always empty there regardless of the recording fix. Wired up, and the LiteAgent path now reads from whichever agent the executor was handed (`original_agent` under kickoff, `self` standalone) rather than assuming. `last_tool_failures` returns a copy, so a caller cannot mutate the agent's record or watch it shift mid-run. Testing: 7 further tests, 52 total, covering the timeout wrapper, the retry limit, kickoff reset, the kickoff output path, copy semantics and guardrail accumulation. Full suite matches baseline exactly at 377 pre-existing failures; mypy clean on every changed file. Co-Authored-By: Claude Opus 5 (1M context) <noreply@anthropic.com> Claude-Session: https://claude.ai/code/session_01ETacm2dMASfpMAYUiDu5YG