mirror of
https://github.com/crewAIInc/crewAI.git
synced 2026-08-13 09:48:03 +00:00
Some checks failed
CodeQL Advanced / Analyze (actions) (push) Has been cancelled
CodeQL Advanced / Analyze (python) (push) Has been cancelled
Check Documentation Broken Links / Check broken links (push) Has been cancelled
Vulnerability Scan / Detect changes (push) Has been cancelled
Vulnerability Scan / pip-audit (push) Has been cancelled
Nightly Canary Release / Check for new commits (push) Has been cancelled
Nightly Canary Release / Build nightly packages (push) Has been cancelled
Nightly Canary Release / Publish nightly to PyPI (push) Has been cancelled
* feat(flow): report flow outcome and human-in-the-loop signals A flow reported only that it started. FlowFinishedEvent, FlowFailedEvent, MethodExecutionFailedEvent, MethodExecutionPausedEvent and FlowPausedEvent all reached the console formatter and stopped there, and FlowInputRequestedEvent, FlowInputReceivedEvent and ConversationTurnFailedEvent had no listener at all - so success rate, failure rate and every HITL pause were unmeasurable. Adds flow:completed, flow:failed, flow:method_failed, flow:paused, flow:hitl_paused, flow:input_requested, flow:input_received and flow:conversation_turn_failed as feature-usage spans, which the existing feature-usage aggregation already reads. Deliberately does not hold the Flow Execution span open to measure duration: flow_executions_daily_target counts those spans at start, so a run that never finishes would disappear from the count entirely. Duration needs its own span. Counts only - flow names, method names, error text and flow state are never recorded. Co-Authored-By: Claude Opus 5 (1M context) <noreply@anthropic.com> Claude-Session: https://claude.ai/code/session_01ASfWmW3RGy4qAQm6s8U9jH * feat(flow): record how long a flow ran Adds a Flow Completed span carrying flow_name, duration_ms and outcome, emitted when a flow finishes or fails. Elapsed time comes from a monotonic stamp taken at flow start and cleared on use. Kept separate from the Flow Execution span rather than holding that one open: it is emitted and closed at start and the daily aggregate counts it, so holding it would drop every run that is killed or crashes from the execution count. A killed run now simply has no Flow Completed row, and the count is unaffected. Elapsed time is an explicit duration_ms attribute rather than the span's own duration, which the ingestion pipeline stores as a suffixed string ("0.0000184s") that downstream aggregation parses to zero. Co-Authored-By: Claude Opus 5 (1M context) <noreply@anthropic.com> Claude-Session: https://claude.ai/code/session_01ASfWmW3RGy4qAQm6s8U9jH * feat(flow): tag flow origin and report resumed runs Two gaps found while testing the pause/resume path end to end. Resumed runs were invisible. There is no resume event: a restored run re-enters through kickoff(), so it looked identical to a fresh start. flow:resumed is derived from _is_execution_resuming at flow start, which makes flow:paused - flow:resumed the abandonment rate. Flow counts are dominated by CrewAI's own AgentExecutor, which is itself a Flow and runs once per agent execution - it is the top flow in the warehouse by a wide margin. Nothing distinguished it from a user's flows except guessing at the name. Both Flow Execution and Flow Completed now carry origin: "internal" when the flow class is defined under crewai.*, "user" otherwise. Tagging only the new span would have left the existing daily count unsplittable. Both span methods take origin with a default, so their signatures stay backward compatible. Co-Authored-By: Claude Opus 5 (1M context) <noreply@anthropic.com> Claude-Session: https://claude.ai/code/session_01ASfWmW3RGy4qAQm6s8U9jH * fix(flow): scope outcome and resume signals to user flows Two findings from review, both confirmed against the code. Outcome features counted CrewAI's own flows. The agent executor, memory encoding and memory recall are all Flows and all set suppress_flow_events; they run far more often than anything a user wrote, so flow:completed, flow:failed and flow:method_failed were mostly bookkeeping. Those three are now emitted only for flows the caller wrote. Internal outcomes are still recorded on the Flow Completed span, which carries origin. flow:resumed counted checkpoint restores. _is_execution_resuming is set both by from_pending (a human pause) and by a checkpoint restore that never paused for anyone, so resumes could exceed pauses and the abandonment rate was unusable. Keyed off _pending_feedback_context instead, which only from_pending sets. Co-Authored-By: Claude Opus 5 (1M context) <noreply@anthropic.com> Claude-Session: https://claude.ai/code/session_01ASfWmW3RGy4qAQm6s8U9jH * fix(flow): declare internal flows instead of inferring them Three findings from review, all confirmed against the code. Gating on suppress_flow_events was wrong. That flag asks for console quiet and is a public field, so a caller who set it on their own flow silently lost flow:completed, flow:failed and flow:method_failed. Deciding origin from the defining module was also wrong. Flow.from_declaration() returns a Flow typed in crewai.flow.flow, so a caller's declarative flow was reported as one of CrewAI's own - the inversion this split exists to prevent. Both had the same root cause: the discriminator was inferred. Flow now declares is_crewai_internal, set on the agent executor and the memory encoding/recall flows, and one helper serves both origin and the outcome gate. A failed conversational session was reported as completed. Its session closes with FlowFinishedEvent whatever happened, so a failed turn produced flow:conversation_turn_failed and flow:completed together. The turn failure is now recorded on the flow and read back when the session finishes. Co-Authored-By: Claude Opus 5 (1M context) <noreply@anthropic.com> Claude-Session: https://claude.ai/code/session_01ASfWmW3RGy4qAQm6s8U9jH * refactor(flow): report flow lifecycle as spans, not feature usage Flow start, completion, pause and method failure are lifecycle facts, and the lifecycle is reported as spans everywhere else. Reporting them through feature usage put them in a table that aggregates on the feature string alone - it cannot carry origin, duration or outcome, so those signals could never be split between a user's flows and the ones CrewAI runs for itself. Adds Flow Paused and Flow Method Failed spans, and a resumed marker on Flow Execution so a run restored from a pause is not counted as a second fresh start. Removes the duplicate feature rows for completed, failed, method_failed, paused and resumed - every one of those facts is now on a span, with more attached to it than the feature row ever carried. Feature usage keeps only genuine adoption signals: flow:hitl_paused, flow:input_requested, flow:input_received and flow:conversation_turn_failed. Also clears the conversational turn-failure flag on every terminal path. A turn that failed without deferred finalization ends via FlowFailedEvent, and the flag left set there marked the next run on that instance as failed. Co-Authored-By: Claude Opus 5 (1M context) <noreply@anthropic.com> Claude-Session: https://claude.ai/code/session_01ASfWmW3RGy4qAQm6s8U9jH * test(flow): update the flow_execution_span caller for the resumed argument Adding the resumed marker changed a signature that tests/utilities/test_events.py asserts on exactly, and that assertion was not re-run before pushing. Co-Authored-By: Claude Opus 5 (1M context) <noreply@anthropic.com> Claude-Session: https://claude.ai/code/session_01ASfWmW3RGy4qAQm6s8U9jH * test(flow): make the checkpoint-restore guard actually guard The test asserted that flow:resumed was absent from feature usage, but that signal moved onto the Flow Execution span. The assertion could no longer fail, so a regression that mis-tagged checkpoint restores as resumes would have gone unnoticed. Now asserts the resumed attribute, and waits for the handlers: the manual emit dispatches asynchronously, so the previous shape also read its result before the listener had run. Confirmed it discriminates - keying resumed off _is_execution_resuming again fails it with [('RestoredFlow', True)] == [('RestoredFlow', False)]. Co-Authored-By: Claude Opus 5 (1M context) <noreply@anthropic.com> Claude-Session: https://claude.ai/code/session_01ASfWmW3RGy4qAQm6s8U9jH * fix(telemetry): record the resumed marker as a string Verified end to end against the live collector and ClickHouse: the pipeline encodes a boolean attribute as the presence of a vBool key, so false arrives as the key simply being absent. That is invisible in the schema and easy to read wrongly - crew_memory is extracted as "the attribute exists" and consequently reports 1 for 99.8% of crews against a field that defaults to False. A string leaves nothing to infer. Confirmed in the warehouse: the emitted span reads resumed = "false". Adds direct coverage for the attributes each flow span records, including both resumed values, and resets the Telemetry singleton in the helper so more than one span method can be exercised per session. Co-Authored-By: Claude Opus 5 (1M context) <noreply@anthropic.com> Claude-Session: https://claude.ai/code/session_01ASfWmW3RGy4qAQm6s8U9jH --------- Co-authored-by: Claude Opus 5 (1M context) <noreply@anthropic.com>
618 lines
19 KiB
Python
618 lines
19 KiB
Python
"""Flow outcome and human-in-the-loop signals must reach telemetry.
|
|
|
|
Driven through real ``Flow`` executions rather than by emitting events directly,
|
|
so these fail if the event bus, the listener wiring, or the emitting call site
|
|
changes - not just if the listener body does.
|
|
|
|
Before this, a flow reported only that it *started*: ``FlowFinishedEvent``,
|
|
``FlowFailedEvent``, ``MethodExecutionFailedEvent``, ``MethodExecutionPausedEvent``
|
|
and ``FlowPausedEvent`` all reached the console formatter and stopped there, and
|
|
the input and conversation-failure events had no listener at all.
|
|
"""
|
|
|
|
from __future__ import annotations
|
|
|
|
import contextlib
|
|
import time
|
|
|
|
import pytest
|
|
|
|
from crewai.flow.async_feedback import HumanFeedbackPending, PendingFeedbackContext
|
|
from crewai.flow.flow import Flow, listen, start
|
|
from crewai.flow.human_feedback import human_feedback
|
|
from crewai.flow.input_provider import InputResponse
|
|
|
|
from ..utils import wait_for_event_handlers
|
|
|
|
|
|
def _reregister_listener() -> None:
|
|
"""Re-subscribe the global listener to the event bus.
|
|
|
|
The repo-wide ``cleanup_event_handlers`` fixture clears every handler after
|
|
each test, so anything relying on the shared listener sees an empty bus
|
|
unless it happens to run first.
|
|
"""
|
|
from crewai.events import event_listener as listener_module
|
|
from crewai.events.event_bus import crewai_event_bus
|
|
from crewai.events.types.flow_events import FlowStartedEvent
|
|
|
|
# Only when the bus is empty: subscribing a second time registers a fresh
|
|
# set of closures, and every handler then fires twice.
|
|
if crewai_event_bus._sync_handlers.get(FlowStartedEvent):
|
|
return
|
|
|
|
listener_module.event_listener.setup_listeners(crewai_event_bus)
|
|
|
|
|
|
@pytest.fixture
|
|
def flow_spans(monkeypatch: pytest.MonkeyPatch) -> list[tuple[str, str]]:
|
|
"""Record (flow_name, origin) for every Flow Execution span."""
|
|
from crewai.events import event_listener as listener_module
|
|
|
|
_reregister_listener()
|
|
|
|
recorded: list[tuple[str, str]] = []
|
|
monkeypatch.setattr(
|
|
listener_module.event_listener._telemetry,
|
|
"flow_execution_span",
|
|
lambda flow_name, node_names, origin="user", resumed=False: recorded.append(
|
|
(flow_name, origin)
|
|
),
|
|
)
|
|
return recorded
|
|
|
|
|
|
@pytest.fixture
|
|
def starts(monkeypatch: pytest.MonkeyPatch) -> list[tuple[str, bool]]:
|
|
"""Record (flow_name, resumed) for every Flow Execution span."""
|
|
from crewai.events import event_listener as listener_module
|
|
|
|
_reregister_listener()
|
|
|
|
recorded: list[tuple[str, bool]] = []
|
|
monkeypatch.setattr(
|
|
listener_module.event_listener._telemetry,
|
|
"flow_execution_span",
|
|
lambda flow_name, node_names, origin="user", resumed=False: recorded.append(
|
|
(flow_name, resumed)
|
|
),
|
|
)
|
|
return recorded
|
|
|
|
|
|
@pytest.fixture
|
|
def pauses(monkeypatch: pytest.MonkeyPatch) -> list[tuple[str, str]]:
|
|
"""Record (flow_name, origin) for every Flow Paused span."""
|
|
from crewai.events import event_listener as listener_module
|
|
|
|
_reregister_listener()
|
|
|
|
recorded: list[tuple[str, str]] = []
|
|
monkeypatch.setattr(
|
|
listener_module.event_listener._telemetry,
|
|
"flow_paused_span",
|
|
lambda flow_name, origin="user": recorded.append((flow_name, origin)),
|
|
)
|
|
return recorded
|
|
|
|
|
|
@pytest.fixture
|
|
def method_failures(monkeypatch: pytest.MonkeyPatch) -> list[tuple[str, str]]:
|
|
"""Record (flow_name, origin) for every Flow Method Failed span."""
|
|
from crewai.events import event_listener as listener_module
|
|
|
|
_reregister_listener()
|
|
|
|
recorded: list[tuple[str, str]] = []
|
|
monkeypatch.setattr(
|
|
listener_module.event_listener._telemetry,
|
|
"flow_method_failed_span",
|
|
lambda flow_name, origin="user": recorded.append((flow_name, origin)),
|
|
)
|
|
return recorded
|
|
|
|
|
|
@pytest.fixture
|
|
def durations(monkeypatch: pytest.MonkeyPatch) -> list[tuple[str, float, str]]:
|
|
"""Record every (flow_name, duration_ms, outcome) the listener reports."""
|
|
from crewai.events import event_listener as listener_module
|
|
|
|
_reregister_listener()
|
|
|
|
recorded: list[tuple[str, float, str]] = []
|
|
monkeypatch.setattr(
|
|
listener_module.event_listener._telemetry,
|
|
"flow_completed_span",
|
|
lambda flow_name, duration_ms, outcome, origin="user": recorded.append(
|
|
(flow_name, duration_ms, outcome)
|
|
),
|
|
)
|
|
return recorded
|
|
|
|
|
|
@pytest.fixture
|
|
def features(monkeypatch: pytest.MonkeyPatch) -> list[str]:
|
|
"""Record every feature the listener reports for a real flow run.
|
|
|
|
Observes the telemetry boundary rather than exported spans: the suite builds
|
|
the Telemetry singleton with collection disabled, so it has no provider to
|
|
export through, and replacing that singleton mid-session leaves the event
|
|
bus without its handlers. That the recorded features become spans is covered
|
|
by ``test_tracer_isolation``.
|
|
"""
|
|
from crewai.events import event_listener as listener_module
|
|
|
|
_reregister_listener()
|
|
|
|
recorded: list[str] = []
|
|
monkeypatch.setattr(
|
|
listener_module.event_listener._telemetry,
|
|
"feature_usage_span",
|
|
recorded.append,
|
|
)
|
|
return recorded
|
|
|
|
|
|
def test_completed_flow_reports_its_outcome(
|
|
durations: list[tuple[str, float, str]],
|
|
) -> None:
|
|
"""Outcome is a lifecycle fact, so it belongs on a span, not a feature."""
|
|
|
|
class OkFlow(Flow):
|
|
@start()
|
|
def go(self) -> str:
|
|
return "ok"
|
|
|
|
OkFlow().kickoff()
|
|
|
|
assert [(n, o) for n, _d, o in durations] == [("OkFlow", "completed")]
|
|
|
|
|
|
def test_failed_flow_reports_the_failure_and_the_method(
|
|
durations: list[tuple[str, float, str]], method_failures: list[tuple[str, str]]
|
|
) -> None:
|
|
class BoomFlow(Flow):
|
|
@start()
|
|
def go(self) -> str:
|
|
raise RuntimeError("boom")
|
|
|
|
with pytest.raises(RuntimeError, match="boom"):
|
|
BoomFlow().kickoff()
|
|
|
|
assert [(n, o) for n, _d, o in durations] == [("BoomFlow", "failed")]
|
|
assert ("BoomFlow", "user") in method_failures
|
|
|
|
|
|
def test_a_failed_flow_is_still_counted_as_an_execution(
|
|
monkeypatch: pytest.MonkeyPatch,
|
|
) -> None:
|
|
"""The start-time span must survive, or aborted runs vanish from counts.
|
|
|
|
``flow_executions_daily_target`` counts ``Flow Execution`` spans, emitted
|
|
when the flow starts. Holding that span open until completion to measure
|
|
duration - the obvious way to add duration - would drop every run that never
|
|
finishes, so the outcome signals are reported separately instead.
|
|
"""
|
|
from crewai.events import event_listener as listener_module
|
|
|
|
_reregister_listener()
|
|
|
|
started: list[str] = []
|
|
monkeypatch.setattr(
|
|
listener_module.event_listener._telemetry,
|
|
"flow_execution_span",
|
|
lambda flow_name, node_names, origin="user", resumed=False: started.append(
|
|
flow_name
|
|
),
|
|
)
|
|
|
|
class BoomFlow(Flow):
|
|
@start()
|
|
def go(self) -> str:
|
|
raise RuntimeError("boom")
|
|
|
|
with pytest.raises(RuntimeError, match="boom"):
|
|
BoomFlow().kickoff()
|
|
|
|
assert "BoomFlow" in started
|
|
|
|
|
|
def test_requesting_input_reports_both_sides(features: list[str]) -> None:
|
|
class StubProvider:
|
|
def request_input(self, message: str, flow: Flow, metadata=None):
|
|
return InputResponse(value="typed answer")
|
|
|
|
class AskFlow(Flow):
|
|
@start()
|
|
def go(self) -> str:
|
|
return self.ask("What topic?")
|
|
|
|
AskFlow(input_provider=StubProvider()).kickoff()
|
|
|
|
emitted = features
|
|
assert "flow:input_requested" in emitted
|
|
assert "flow:input_received" in emitted
|
|
|
|
|
|
def test_paused_flow_reports_the_pause(
|
|
features: list[str], pauses: list[tuple[str, str]]
|
|
) -> None:
|
|
"""An async feedback provider pauses the flow; both signals must land."""
|
|
|
|
class AsyncProvider:
|
|
def request_feedback(self, context: PendingFeedbackContext, flow: Flow) -> str:
|
|
raise HumanFeedbackPending(context=context)
|
|
|
|
class PausingFlow(Flow):
|
|
@start()
|
|
@human_feedback(message="Review:", provider=AsyncProvider())
|
|
def generate(self) -> str:
|
|
return "content"
|
|
|
|
@listen(generate)
|
|
def process(self, result) -> str:
|
|
return f"processed: {result.feedback}"
|
|
|
|
# Whether the pause surfaces as an exception depends on the persistence
|
|
# backend in use; the signals must land either way.
|
|
with contextlib.suppress(BaseException):
|
|
PausingFlow().kickoff()
|
|
|
|
# The pause itself is lifecycle and lands on a span; that a human-feedback
|
|
# method was what paused is genuine feature adoption.
|
|
assert ("PausingFlow", "user") in pauses
|
|
assert "flow:hitl_paused" in features
|
|
|
|
|
|
def test_failed_conversation_turn_is_reported(features: list[str]) -> None:
|
|
"""Only completed turns were tracked, so failure rate was unknowable."""
|
|
|
|
class FailingChat(Flow):
|
|
conversational = True
|
|
|
|
@start()
|
|
def begin(self) -> str:
|
|
raise RuntimeError("turn exploded")
|
|
|
|
with pytest.raises(RuntimeError, match="turn exploded"):
|
|
FailingChat().handle_turn("hello")
|
|
|
|
assert "flow:conversation_turn_failed" in features
|
|
|
|
|
|
def test_no_method_names_or_error_text_are_recorded(
|
|
method_failures: list[tuple[str, str]],
|
|
durations: list[tuple[str, float, str]],
|
|
features: list[str],
|
|
) -> None:
|
|
"""Method names and error text are user-authored and must not be sent.
|
|
|
|
The flow name is recorded, as it already is for flow creation and
|
|
execution, so it is deliberately not asserted against here.
|
|
"""
|
|
|
|
class SecretNamedFlow(Flow):
|
|
@start()
|
|
def my_secret_method_name(self) -> str:
|
|
raise RuntimeError("secret error detail")
|
|
|
|
with pytest.raises(RuntimeError, match="secret error detail"):
|
|
SecretNamedFlow().kickoff()
|
|
|
|
assert method_failures, "the failure must still be reported"
|
|
recorded = [
|
|
str(value)
|
|
for row in (*method_failures, *durations)
|
|
for value in row
|
|
] + features
|
|
for value in recorded:
|
|
assert "my_secret_method_name" not in value
|
|
assert "secret error detail" not in value
|
|
|
|
|
|
def test_completed_flow_reports_a_real_duration(
|
|
durations: list[tuple[str, float, str]],
|
|
) -> None:
|
|
"""Elapsed time must be measured, not merely present."""
|
|
|
|
class SlowFlow(Flow):
|
|
@start()
|
|
def go(self) -> str:
|
|
time.sleep(0.05)
|
|
return "ok"
|
|
|
|
SlowFlow().kickoff()
|
|
|
|
assert len(durations) == 1
|
|
flow_name, duration_ms, outcome = durations[0]
|
|
assert flow_name == "SlowFlow"
|
|
assert outcome == "completed"
|
|
assert duration_ms >= 50
|
|
|
|
|
|
def test_failed_flow_reports_its_duration_and_outcome(
|
|
durations: list[tuple[str, float, str]],
|
|
) -> None:
|
|
class SlowBoomFlow(Flow):
|
|
@start()
|
|
def go(self) -> str:
|
|
time.sleep(0.05)
|
|
raise RuntimeError("boom")
|
|
|
|
with pytest.raises(RuntimeError, match="boom"):
|
|
SlowBoomFlow().kickoff()
|
|
|
|
assert len(durations) == 1
|
|
flow_name, duration_ms, outcome = durations[0]
|
|
assert flow_name == "SlowBoomFlow"
|
|
assert outcome == "failed"
|
|
assert duration_ms >= 50
|
|
|
|
|
|
def test_no_duration_is_reported_without_a_recorded_start(
|
|
durations: list[tuple[str, float, str]],
|
|
) -> None:
|
|
"""A completion with no observed start reports nothing, and does not raise.
|
|
|
|
A conversational turn can re-emit completion for a restored run, so the
|
|
stamp is genuinely absent rather than impossible.
|
|
"""
|
|
from crewai.events.event_bus import crewai_event_bus
|
|
from crewai.events.types.flow_events import FlowFinishedEvent
|
|
|
|
class NeverStartedFlow(Flow):
|
|
@start()
|
|
def go(self) -> str:
|
|
return "ok"
|
|
|
|
flow = NeverStartedFlow()
|
|
crewai_event_bus.emit(
|
|
flow,
|
|
FlowFinishedEvent(flow_name="NeverStartedFlow", result="ok", state={}),
|
|
)
|
|
|
|
assert durations == []
|
|
|
|
|
|
def test_duration_is_reported_once_per_run(
|
|
durations: list[tuple[str, float, str]],
|
|
) -> None:
|
|
"""The stamp is cleared on use, so a repeated completion cannot double-count."""
|
|
from crewai.events.event_bus import crewai_event_bus
|
|
from crewai.events.types.flow_events import FlowFinishedEvent
|
|
|
|
class OkFlow(Flow):
|
|
@start()
|
|
def go(self) -> str:
|
|
return "ok"
|
|
|
|
flow = OkFlow()
|
|
flow.kickoff()
|
|
crewai_event_bus.emit(
|
|
flow, FlowFinishedEvent(flow_name="OkFlow", result="ok", state={})
|
|
)
|
|
|
|
assert len(durations) == 1
|
|
|
|
|
|
def test_user_authored_flows_are_tagged_as_user(flow_spans) -> None:
|
|
class MyOwnFlow(Flow):
|
|
@start()
|
|
def go(self) -> str:
|
|
return "ok"
|
|
|
|
MyOwnFlow().kickoff()
|
|
|
|
assert ("MyOwnFlow", "user") in flow_spans
|
|
|
|
|
|
def test_crewais_own_agent_executor_is_tagged_internal(flow_spans) -> None:
|
|
"""The agent executor is a Flow and runs once per agent execution.
|
|
|
|
Without an origin tag it is indistinguishable from a user's flows in the
|
|
daily counts, and it dominates them.
|
|
"""
|
|
from crewai import Agent, Crew, Task
|
|
from crewai.llms.base_llm import BaseLLM
|
|
|
|
class StubLLM(BaseLLM):
|
|
def __init__(self) -> None:
|
|
super().__init__(model="stub-model")
|
|
|
|
def call(self, messages, **kwargs) -> str:
|
|
return "Final Answer: done"
|
|
|
|
def supports_function_calling(self) -> bool:
|
|
return False
|
|
|
|
def supports_stop_words(self) -> bool:
|
|
return False
|
|
|
|
def get_context_window_size(self) -> int:
|
|
return 8192
|
|
|
|
agent = Agent(role="R", goal="G", backstory="B", llm=StubLLM())
|
|
task = Task(description="Do it", expected_output="A result", agent=agent)
|
|
Crew(agents=[agent], tasks=[task]).kickoff()
|
|
|
|
origins = {name: origin for name, origin in flow_spans}
|
|
assert origins.get("AgentExecutor") == "internal"
|
|
|
|
|
|
def test_resumed_flow_is_reported(
|
|
tmp_path,
|
|
pauses: list[tuple[str, str]],
|
|
starts: list[tuple[str, bool]],
|
|
durations: list[tuple[str, float, str]],
|
|
) -> None:
|
|
"""A restored run is only visible here - there is no resume event.
|
|
|
|
Without it, a paused flow that was abandoned cannot be told apart from one
|
|
the user came back to.
|
|
"""
|
|
from pydantic import BaseModel
|
|
|
|
from crewai.events.event_bus import crewai_event_bus
|
|
from crewai.events.types.flow_events import FlowPausedEvent
|
|
from crewai.flow.persistence.sqlite import SQLiteFlowPersistence
|
|
|
|
persistence = SQLiteFlowPersistence(str(tmp_path / "flows.db"))
|
|
|
|
class State(BaseModel):
|
|
id: str = "resume-test-1"
|
|
|
|
class AsyncProvider:
|
|
def request_feedback(self, context: PendingFeedbackContext, flow: Flow) -> str:
|
|
raise HumanFeedbackPending(context=context)
|
|
|
|
class ReviewFlow(Flow[State]):
|
|
@start()
|
|
@human_feedback(message="Review:", provider=AsyncProvider())
|
|
def draft(self) -> str:
|
|
return "draft"
|
|
|
|
@listen(draft)
|
|
def finish(self, result) -> str:
|
|
return f"final: {result.feedback}"
|
|
|
|
paused: dict[str, str] = {}
|
|
|
|
@crewai_event_bus.on(FlowPausedEvent)
|
|
def _capture(source, event) -> None:
|
|
paused["flow_id"] = event.flow_id
|
|
|
|
with contextlib.suppress(BaseException):
|
|
ReviewFlow(persistence=persistence).kickoff()
|
|
|
|
assert ("ReviewFlow", "user") in pauses
|
|
assert starts == [("ReviewFlow", False)]
|
|
|
|
flow = ReviewFlow.from_pending(paused["flow_id"], persistence)
|
|
flow.resume("looks good")
|
|
|
|
assert ("ReviewFlow", True) in starts
|
|
assert ("ReviewFlow", "completed") in [(n, o) for n, _d, o in durations]
|
|
|
|
|
|
def test_a_user_flow_that_suppresses_console_events_still_reports(
|
|
durations: list[tuple[str, float, str]],
|
|
) -> None:
|
|
"""``suppress_flow_events`` asks for console quiet, not for no telemetry."""
|
|
|
|
class QuietFlow(Flow):
|
|
suppress_flow_events: bool = True
|
|
|
|
@start()
|
|
def go(self) -> str:
|
|
return "ok"
|
|
|
|
QuietFlow().kickoff()
|
|
|
|
assert [(n, o) for n, _d, o in durations] == [("QuietFlow", "completed")]
|
|
|
|
|
|
def test_a_declarative_flow_is_not_treated_as_internal(
|
|
flow_spans: list[tuple[str, str]],
|
|
) -> None:
|
|
"""``Flow.from_declaration()`` yields a ``Flow``, defined inside crewai.
|
|
|
|
Deciding origin from the defining module would report a caller's
|
|
declarative flow as one of CrewAI's own.
|
|
"""
|
|
flow = Flow.from_declaration(contents={"name": "MyDeclarativeFlow"})
|
|
|
|
assert getattr(type(flow), "is_crewai_internal", False) is False
|
|
|
|
|
|
def test_a_failed_conversation_session_is_not_reported_completed(
|
|
features: list[str], durations: list[tuple[str, float, str]]
|
|
) -> None:
|
|
"""A conversational session closes with FlowFinishedEvent either way.
|
|
|
|
Reading that event at face value counted a failed session as a success,
|
|
alongside the turn-failure signal.
|
|
"""
|
|
|
|
class FailingChat(Flow):
|
|
conversational = True
|
|
|
|
@start()
|
|
def begin(self) -> str:
|
|
raise RuntimeError("turn exploded")
|
|
|
|
chat = FailingChat()
|
|
with pytest.raises(RuntimeError, match="turn exploded"):
|
|
chat.handle_turn("hello")
|
|
chat.finalize_session_traces()
|
|
|
|
assert "flow:conversation_turn_failed" in features
|
|
assert all(outcome != "completed" for _n, _d, outcome in durations)
|
|
|
|
|
|
def test_infrastructure_flows_do_not_pollute_outcome_signals(
|
|
features: list[str], durations: list[tuple[str, float, str]]
|
|
) -> None:
|
|
"""CrewAI's own flows must not be counted as user flow outcomes.
|
|
|
|
The agent executor, memory encoding and memory recall are all Flows and run
|
|
far more often than anything a user wrote. Counting their outcomes in the
|
|
same feature would make ``flow:completed`` mostly bookkeeping. Their outcome
|
|
is still recorded on the Flow Completed span, which carries ``origin``.
|
|
"""
|
|
from crewai import Agent, Crew, Task
|
|
from crewai.llms.base_llm import BaseLLM
|
|
|
|
class StubLLM(BaseLLM):
|
|
def __init__(self) -> None:
|
|
super().__init__(model="stub-model")
|
|
|
|
def call(self, messages, **kwargs) -> str:
|
|
return "Final Answer: done"
|
|
|
|
def supports_function_calling(self) -> bool:
|
|
return False
|
|
|
|
def supports_stop_words(self) -> bool:
|
|
return False
|
|
|
|
def get_context_window_size(self) -> int:
|
|
return 8192
|
|
|
|
agent = Agent(role="R", goal="G", backstory="B", llm=StubLLM())
|
|
task = Task(description="Do it", expected_output="A result", agent=agent)
|
|
Crew(agents=[agent], tasks=[task]).kickoff()
|
|
|
|
# Internal outcomes are still recorded - on the span, tagged internal -
|
|
# they simply do not masquerade as a user's flow finishing.
|
|
assert ("AgentExecutor", "completed") in [
|
|
(name, outcome) for name, _duration, outcome in durations
|
|
]
|
|
assert "flow:completed" not in features
|
|
|
|
|
|
def test_a_checkpoint_restore_is_not_counted_as_a_resume(
|
|
starts: list[tuple[str, bool]],
|
|
) -> None:
|
|
"""Only a run restored from a human pause is marked resumed.
|
|
|
|
``_is_execution_resuming`` is also set by checkpoint restores that never
|
|
paused for anyone. Counting those would push resumes above pauses and make
|
|
the abandonment rate meaningless.
|
|
"""
|
|
from crewai.events.event_bus import crewai_event_bus
|
|
from crewai.events.types.flow_events import FlowStartedEvent
|
|
|
|
class RestoredFlow(Flow):
|
|
@start()
|
|
def go(self) -> str:
|
|
return "ok"
|
|
|
|
flow = RestoredFlow()
|
|
flow._is_execution_resuming = True
|
|
assert flow._pending_feedback_context is None
|
|
|
|
crewai_event_bus.emit(flow, FlowStartedEvent(flow_name="RestoredFlow"))
|
|
wait_for_event_handlers()
|
|
|
|
assert starts == [("RestoredFlow", False)]
|