mirror of
https://github.com/crewAIInc/crewAI.git
synced 2026-09-22 10:56:50 +00:00
feat(cli): crewai eval evaluates the last traced run through AMP (#7649)
* feat(cli): crewai eval evaluates the last traced run through AMP `crewai eval` reads `.crewai/last_run.json` — the record crewAI writes when a traced run's spans reach Wharf — and asks AMP to evaluate that run: POST /crewai_plus/api/v1/tracing/evaluations with the execution id, sending the saved `crewai login` when there is one and nothing otherwise. AMP answers with an evaluation id and a URL; the command prints the URL, opens it, waits for the verdict and prints it (goal gate and the four grades), exit 1 only when the evaluation itself failed. `--run EXECUTION_ID` evaluates another run. With no traced run recorded it offers to turn tracing on for the project (`CREWAI_TRACING_ENABLED=true` in .env, set_key so nothing else in the file moves) and run the crew now with `crewai run`; without a terminal, or declined, it prints the three steps instead. A run that leaves no record behind is explained, never guessed at. AMP's refusals are printed in its own words: a run that needs an account (401 account_required), a refused credential (then `crewai login`), a run AMP does not hold (404), rate limiting (429 with Retry-After), and any other status with AMP's message. The two AMP calls live on the CLI's PlusAPI subclass, so no crewai-core release is needed. Who may evaluate what is AMP's decision, not the command's: an anonymous run once without an account, then it needs one; a run traced while logged in for that organization's members; a deployment execution for members who may see its traces. Co-Authored-By: Claude Fable 5.1 <noreply@anthropic.com> * fix(cli): crewai eval — guard the project, survive AMP blips, name the right subject, read the whole record Review findings, each reproduced before the fix: - The run-it-now offer wrote CREWAI_TRACING_ENABLED into ./.env before checking the directory is a crewAI project, then run_crew() died on a missing pyproject.toml with a traceback. Now: no pyproject.toml → one sentence, exit 1, nothing written. - httpx errors (AMP unreachable, a timeout) surfaced as tracebacks. Now the start says "Could not reach AMP to start the evaluation: …"; while waiting, an unreachable AMP or a 5xx is retried up to POLL_RETRIES consecutive times, then reported with the URL — the evaluation keeps running server-side either way. - The POST now carries a 120 s timeout (AMP reads the run's spans inside it), the poll 30 s. - A 200 whose body has no known status (a non-dict, no status, a status outside queued/running/done/failed) polled forever. Now it stops with the status it saw and the URL. - A 404 without a JSON message read "AMP holds no run <evaluation id>" while polling. _refused takes "run <id>" / "evaluation <id>" and says "AMP answered 404 for <subject>" — AMP's own message still wins. - The record's amp_base_url was ignored; the CLI now evaluates the run at the AMP it was traced to. --run keeps the configured AMP. - The post-run explanation names the third cause: a crewai older than the version that records the last run. Tests for each, plus the previously untested paths: Ctrl-C exits 130, a 2xx without an id, a refusal mid-poll, DMN opens no browser, --run skips the offer. 51 passed in lib/cli/tests (eval + plus_api). Co-Authored-By: Claude Fable 5.1 <noreply@anthropic.com> * fix(cli): crewai eval — the credential goes only to the configured AMP; an explicit yes before running the crew; a done answer needs a well-formed verdict Review findings (CodeRabbit, the PR gate): - The record's amp_base_url was handed to PlusAPI beside the saved login, so a modified .crewai/last_run.json could send the token to any origin. The client is now built from the configured AMP only (CREWAI_PLUS_URL, the saved settings, app.crewai.com); the project's .env is loaded first, as `crewai run` loads it, so the configured AMP is the one the run was traced to. A record naming another address gets a one-line note and no credential. - The offer to turn tracing on and run the crew defaults to no and says the .env change stays; Enter no longer spends a crew run. - A `done` answer whose verdict is missing or malformed (no gate, grades not an object, a grade not an int or null) is a protocol error, exit 1, instead of an INCONCLUSIVE line with exit 0 or an AttributeError. Tests for each: a foreign origin in the record with a saved token, the .env-loaded same-AMP case, seven malformed verdicts, the prompt's text and default. 59 passed (eval + plus_api); ruff and mypy clean. Co-Authored-By: Claude Fable 5.1 <noreply@anthropic.com> * fix(cli): crewai eval — Enter accepts the offer to run the crew (y/n, Y default; the prompt names both effects) João's call (2026-09-20): the confirm is y/n with Y as the default. The prompt still says tracing stays on in .env and that the crew runs now. Co-Authored-By: Claude Fable 5.1 <noreply@anthropic.com> --------- Co-authored-by: Claude Fable 5.1 <noreply@anthropic.com>
This commit is contained in:
505
lib/cli/tests/test_eval_crew.py
Normal file
505
lib/cli/tests/test_eval_crew.py
Normal file
@@ -0,0 +1,505 @@
|
||||
"""`crewai eval`: the last traced run, evaluated through AMP."""
|
||||
|
||||
from __future__ import annotations
|
||||
|
||||
import json
|
||||
import os
|
||||
from pathlib import Path
|
||||
from types import SimpleNamespace
|
||||
|
||||
from click.testing import CliRunner
|
||||
import httpx
|
||||
import pytest
|
||||
from rich.console import Console
|
||||
|
||||
from crewai_cli import eval_crew as eval_module
|
||||
from crewai_cli.cli import eval_command
|
||||
|
||||
|
||||
EXECUTION_ID = "6f31fe1a-20bd-4bfe-a011-25d6b9341f62"
|
||||
URL = "https://evolve.crewai.test/e/ev-1"
|
||||
|
||||
|
||||
class FakeAMP:
|
||||
"""A PlusAPI double: scripted answers, calls recorded."""
|
||||
|
||||
def __init__(self, create=None, statuses=None):
|
||||
self.create = create if create is not None else httpx.Response(
|
||||
202, json={"id": "ev-1", "url": URL, "status": "queued"}
|
||||
)
|
||||
self.statuses = list(statuses or [])
|
||||
self.calls: list[tuple] = []
|
||||
self.api_key = None
|
||||
|
||||
def create_evaluation(self, execution_id):
|
||||
self.calls.append(("create", execution_id))
|
||||
return self.create
|
||||
|
||||
def get_evaluation(self, evaluation_id):
|
||||
self.calls.append(("get", evaluation_id))
|
||||
return self.statuses.pop(0) if self.statuses else httpx.Response(200, json={"id": evaluation_id, "status": "running"})
|
||||
|
||||
|
||||
def done(gate="passed", grades=None):
|
||||
return httpx.Response(200, json={
|
||||
"id": "ev-1", "status": "done", "url": URL,
|
||||
"verdict": {"gate": gate, "grades": grades if grades is not None else {"goal": 5, "quality": 4, "process": 5, "cost": None}},
|
||||
})
|
||||
|
||||
|
||||
@pytest.fixture
|
||||
def project(tmp_path, monkeypatch):
|
||||
monkeypatch.chdir(tmp_path)
|
||||
monkeypatch.setattr(eval_module, "console", Console(width=240)) # one sentence per line in the captured output
|
||||
monkeypatch.setattr(eval_module, "get_or_create_project_id", lambda: None)
|
||||
monkeypatch.setattr(eval_module, "saved_login", lambda: "login-token")
|
||||
monkeypatch.setattr(eval_module.time, "sleep", lambda seconds: None)
|
||||
monkeypatch.setattr(eval_module, "is_dmn_mode_enabled", lambda: False)
|
||||
# This machine is logged in to https://amp.test (`crewai enterprise configure`).
|
||||
monkeypatch.setattr(eval_module, "Settings", lambda: SimpleNamespace(enterprise_base_url="https://amp.test"))
|
||||
monkeypatch.delenv("CREWAI_PLUS_URL", raising=False)
|
||||
opened: list[str] = []
|
||||
monkeypatch.setattr(eval_module.webbrowser, "open", lambda url: opened.append(url) or True)
|
||||
return tmp_path, opened
|
||||
|
||||
|
||||
def record_last_run(directory: Path, execution_id: str = EXECUTION_ID, **fields) -> None:
|
||||
(directory / ".crewai").mkdir(exist_ok=True)
|
||||
record = {"execution_id": execution_id, "tier": "ephemeral", "amp_base_url": "https://amp.test", **fields}
|
||||
(directory / ".crewai" / "last_run.json").write_text(json.dumps(record))
|
||||
|
||||
|
||||
def install(monkeypatch, amp: FakeAMP, configured_amp: str = "https://amp.test") -> FakeAMP:
|
||||
def build(api_key=None, base_url=None):
|
||||
# PlusAPI's own resolution: explicit, then CREWAI_PLUS_URL, then the saved settings.
|
||||
amp.api_key = api_key
|
||||
amp.base_url = base_url or os.environ.get("CREWAI_PLUS_URL") or configured_amp
|
||||
return amp
|
||||
|
||||
monkeypatch.setattr(eval_module, "PlusAPI", build)
|
||||
return amp
|
||||
|
||||
|
||||
def test_the_last_run_is_evaluated_the_url_opened_and_the_verdict_printed(project, monkeypatch, capsys):
|
||||
directory, opened = project
|
||||
record_last_run(directory)
|
||||
amp = install(monkeypatch, FakeAMP(statuses=[httpx.Response(200, json={"id": "ev-1", "status": "running"}), done()]))
|
||||
|
||||
eval_module.eval_crew()
|
||||
|
||||
out = capsys.readouterr().out
|
||||
assert amp.api_key == "login-token"
|
||||
assert amp.calls == [("create", EXECUTION_ID), ("get", "ev-1"), ("get", "ev-1")]
|
||||
assert opened == [URL]
|
||||
assert EXECUTION_ID in out and URL in out
|
||||
assert "Goal gate: PASSED" in out and "goal 5/5" in out and "cost not measured" in out
|
||||
|
||||
|
||||
def test_run_names_another_execution_and_an_anonymous_caller_sends_no_token(project, monkeypatch, capsys):
|
||||
directory, _ = project
|
||||
record_last_run(directory)
|
||||
monkeypatch.setattr(eval_module, "saved_login", lambda: None)
|
||||
amp = install(monkeypatch, FakeAMP(statuses=[done("failed")]))
|
||||
|
||||
eval_module.eval_crew(run_id="other-run")
|
||||
|
||||
assert amp.api_key is None
|
||||
assert amp.calls[0] == ("create", "other-run")
|
||||
assert "Goal gate: FAILED" in capsys.readouterr().out
|
||||
|
||||
|
||||
def test_a_failed_evaluation_exits_one_with_amps_reason(project, monkeypatch, capsys):
|
||||
directory, _ = project
|
||||
record_last_run(directory)
|
||||
install(monkeypatch, FakeAMP(statuses=[httpx.Response(200, json={"id": "ev-1", "status": "failed", "error": "the judge was unreachable"})]))
|
||||
|
||||
with pytest.raises(SystemExit) as exit_:
|
||||
eval_module.eval_crew()
|
||||
|
||||
assert exit_.value.code == 1
|
||||
assert "Evaluation failed: the judge was unreachable" in capsys.readouterr().out
|
||||
|
||||
|
||||
@pytest.mark.parametrize(
|
||||
("response", "expected"),
|
||||
[
|
||||
(httpx.Response(401, json={"error": "account_required", "message": "Execution x was already read once without an account. Log in with `crewai login`, or create an account, to read it again."}),
|
||||
"already read once without an account"),
|
||||
(httpx.Response(401, json={"error": "bad_credentials", "message": "Bad credentials"}), "Bad credentials. Log in with `crewai login`"),
|
||||
(httpx.Response(404, json={"error": "trace_not_found", "message": "No spans recorded for execution 6f31"}), "No spans recorded for execution 6f31"),
|
||||
(httpx.Response(404, text="<html>Page not found</html>"), f"AMP answered 404 for run {EXECUTION_ID}."),
|
||||
(httpx.Response(202, json={"url": URL}), "AMP answered without an evaluation id (202)."),
|
||||
(httpx.Response(429, json={"error": "rate_limit_exceeded", "message": "Too many requests"}, headers={"Retry-After": "60"}), "Too many requests — retry after 60s"),
|
||||
(httpx.Response(503, json={"error": "service_unavailable", "message": "Wharf could not list the spans"}), "AMP answered 503: Wharf could not list the spans"),
|
||||
(httpx.Response(500, text="boom"), "AMP answered 500."),
|
||||
],
|
||||
)
|
||||
def test_amps_refusals_are_printed_in_its_words_and_exit_one(project, monkeypatch, capsys, response, expected):
|
||||
directory, opened = project
|
||||
record_last_run(directory)
|
||||
install(monkeypatch, FakeAMP(create=response))
|
||||
|
||||
with pytest.raises(SystemExit) as exit_:
|
||||
eval_module.eval_crew()
|
||||
|
||||
assert exit_.value.code == 1
|
||||
assert expected in capsys.readouterr().out
|
||||
assert opened == []
|
||||
|
||||
|
||||
def test_the_credential_goes_only_to_the_configured_amp_never_to_an_address_off_the_record(project, monkeypatch, capsys):
|
||||
directory, _ = project
|
||||
record_last_run(directory, amp_base_url="https://evil.example/steal")
|
||||
amp = install(monkeypatch, FakeAMP(statuses=[done()]), configured_amp="https://app.crewai.com")
|
||||
|
||||
eval_module.eval_crew()
|
||||
|
||||
assert amp.api_key == "login-token" and amp.base_url == "https://app.crewai.com" # PlusAPI got no base_url
|
||||
out = capsys.readouterr().out
|
||||
assert "The run was traced to https://evil.example/steal; evaluating at the configured AMP https://app.crewai.com." in out
|
||||
|
||||
# The project's .env is what `crewai run` traced with, so it is loaded first. This machine is
|
||||
# logged in to that AMP, so the credential goes with the request and nothing is remarked on.
|
||||
(directory / ".env").write_text("CREWAI_PLUS_URL=https://amp.test\n")
|
||||
record_last_run(directory, amp_base_url="https://amp.test/")
|
||||
amp = install(monkeypatch, FakeAMP(statuses=[done()]))
|
||||
eval_module.eval_crew()
|
||||
assert eval_module.os.environ["CREWAI_PLUS_URL"] == "https://amp.test"
|
||||
assert amp.api_key == "login-token" and amp.base_url == "https://amp.test"
|
||||
out = capsys.readouterr().out
|
||||
assert "was traced to" not in out and "Reading anonymously" not in out
|
||||
|
||||
|
||||
def test_a_project_may_point_at_another_amp_but_never_gets_the_saved_login(project, monkeypatch, capsys):
|
||||
"""A .env can send the request elsewhere — that is how a self-hosted project is wired —
|
||||
but the token goes only to an AMP this machine is logged in to."""
|
||||
directory, _ = project
|
||||
(directory / ".env").write_text("CREWAI_PLUS_URL=https://evil.example\n")
|
||||
record_last_run(directory, amp_base_url="https://evil.example")
|
||||
amp = install(monkeypatch, FakeAMP(statuses=[done()]))
|
||||
|
||||
eval_module.eval_crew() # still works: AMP reads it anonymously
|
||||
|
||||
assert amp.base_url == "https://evil.example" # the request follows the project
|
||||
assert amp.api_key is None # the credential does not
|
||||
out = capsys.readouterr().out
|
||||
assert "Reading anonymously: https://evil.example is not an AMP this machine is logged in to." in out
|
||||
assert "crewai enterprise configure" in out
|
||||
|
||||
|
||||
def test_a_trusted_amp_over_plain_http_still_gets_no_credential(project, monkeypatch, capsys):
|
||||
"""A cleartext connection is not a place to put a bearer token, trusted or not."""
|
||||
directory, _ = project
|
||||
monkeypatch.setenv("CREWAI_PLUS_URL", "http://amp.internal") # exported, so it IS trusted
|
||||
record_last_run(directory, amp_base_url="http://amp.internal")
|
||||
amp = install(monkeypatch, FakeAMP(statuses=[done()]))
|
||||
|
||||
eval_module.eval_crew()
|
||||
|
||||
assert amp.base_url == "http://amp.internal" and amp.api_key is None
|
||||
assert "would carry the login over plain HTTP" in capsys.readouterr().out
|
||||
|
||||
|
||||
def test_plain_http_to_this_machine_is_fine_for_local_development(project, monkeypatch, capsys):
|
||||
directory, _ = project
|
||||
monkeypatch.setenv("CREWAI_PLUS_URL", "http://localhost:3000")
|
||||
record_last_run(directory, amp_base_url="http://localhost:3000")
|
||||
amp = install(monkeypatch, FakeAMP(statuses=[done()]))
|
||||
|
||||
eval_module.eval_crew()
|
||||
|
||||
assert amp.api_key == "login-token"
|
||||
assert "Reading anonymously" not in capsys.readouterr().out
|
||||
|
||||
|
||||
@pytest.mark.parametrize(
|
||||
("origin", "encrypted"),
|
||||
[
|
||||
("https://app.crewai.com", True),
|
||||
("http://localhost:3000", True),
|
||||
("http://127.0.0.1:8000", True),
|
||||
("http://[::1]:8000", True),
|
||||
("http://amp.localhost", True),
|
||||
("http://amp.internal", False),
|
||||
("http://169.254.169.254", False),
|
||||
("ftp://amp.test", False),
|
||||
(None, False),
|
||||
],
|
||||
)
|
||||
def test_which_connections_may_carry_the_login(origin, encrypted):
|
||||
assert eval_module._encrypted(origin) is encrypted
|
||||
|
||||
|
||||
def test_an_amp_exported_in_this_shell_is_trusted(project, monkeypatch, capsys):
|
||||
directory, _ = project
|
||||
monkeypatch.setenv("CREWAI_PLUS_URL", "https://shell.amp.test") # exported before the project is read
|
||||
record_last_run(directory, amp_base_url="https://shell.amp.test")
|
||||
amp = install(monkeypatch, FakeAMP(statuses=[done()]))
|
||||
|
||||
eval_module.eval_crew()
|
||||
|
||||
assert amp.base_url == "https://shell.amp.test" and amp.api_key == "login-token"
|
||||
assert "Reading anonymously" not in capsys.readouterr().out
|
||||
|
||||
|
||||
@pytest.mark.parametrize(
|
||||
("url", "origin"),
|
||||
[
|
||||
("https://amp.test", "https://amp.test"),
|
||||
("https://AMP.Test/crewai_plus/", "https://amp.test"),
|
||||
("http://localhost:8000/x", "http://localhost:8000"),
|
||||
("app.crewai.com", None), # no scheme: not an origin, never trusted
|
||||
("", None),
|
||||
(None, None),
|
||||
],
|
||||
)
|
||||
def test_an_origin_is_scheme_and_host_only(url, origin):
|
||||
assert eval_module._origin(url) == origin
|
||||
|
||||
|
||||
@pytest.mark.parametrize(
|
||||
"verdict",
|
||||
[
|
||||
None,
|
||||
"passed",
|
||||
[],
|
||||
{"gate": "passed"},
|
||||
{"gate": None, "grades": {}},
|
||||
{"gate": "passed", "grades": "5/5"},
|
||||
{"gate": "passed", "grades": {"goal": "five"}},
|
||||
{"gate": "passed", "grades": {"goal": True}}, # a bool is an int to Python, never a grade
|
||||
{"gate": "passed", "grades": {"goal": 6}}, # out of the 1..5 range
|
||||
{"gate": "passed", "grades": {"goal": 0}},
|
||||
{"gate": "passed", "grades": {"goal": 4.5}},
|
||||
],
|
||||
)
|
||||
def test_a_done_answer_without_a_well_formed_verdict_is_a_protocol_error(project, monkeypatch, capsys, verdict):
|
||||
directory, _ = project
|
||||
record_last_run(directory)
|
||||
body = {"id": "ev-1", "status": "done", "url": URL}
|
||||
if verdict is not None:
|
||||
body["verdict"] = verdict
|
||||
install(monkeypatch, FakeAMP(statuses=[httpx.Response(200, json=body)]))
|
||||
|
||||
with pytest.raises(SystemExit) as exit_:
|
||||
eval_module.eval_crew()
|
||||
|
||||
assert exit_.value.code == 1
|
||||
assert f"AMP answered done without a verdict (protocol error); follow it at {URL}." in capsys.readouterr().out
|
||||
|
||||
|
||||
def test_amp_unreachable_at_the_start_is_a_sentence_not_a_traceback(project, monkeypatch, capsys):
|
||||
directory, _ = project
|
||||
record_last_run(directory)
|
||||
amp = install(monkeypatch, FakeAMP())
|
||||
monkeypatch.setattr(amp, "create_evaluation", lambda execution_id: (_ for _ in ()).throw(httpx.ConnectError("connection refused")))
|
||||
|
||||
with pytest.raises(SystemExit) as exit_:
|
||||
eval_module.eval_crew()
|
||||
|
||||
assert exit_.value.code == 1
|
||||
assert "Could not reach AMP to start the evaluation: connection refused" in capsys.readouterr().out
|
||||
|
||||
|
||||
def test_while_waiting_a_blip_is_retried_and_a_streak_is_reported(project, monkeypatch, capsys):
|
||||
directory, _ = project
|
||||
record_last_run(directory)
|
||||
amp = install(monkeypatch, FakeAMP(statuses=[httpx.Response(502), httpx.Response(503, json={"error": "service_unavailable", "message": "crew-optimize is down"}), done()]))
|
||||
|
||||
eval_module.eval_crew() # two bad polls, then the verdict
|
||||
|
||||
assert "Goal gate: PASSED" in capsys.readouterr().out
|
||||
assert amp.calls.count(("get", "ev-1")) == 3
|
||||
|
||||
record_last_run(directory)
|
||||
amp = install(monkeypatch, FakeAMP(statuses=[httpx.Response(503, json={"error": "service_unavailable", "message": "crew-optimize is down"})] * eval_module.POLL_RETRIES))
|
||||
with pytest.raises(SystemExit) as exit_:
|
||||
eval_module.eval_crew()
|
||||
assert exit_.value.code == 1
|
||||
assert "AMP answered 503: crew-optimize is down" in capsys.readouterr().out
|
||||
assert amp.calls.count(("get", "ev-1")) == eval_module.POLL_RETRIES
|
||||
|
||||
amp = install(monkeypatch, FakeAMP())
|
||||
monkeypatch.setattr(amp, "get_evaluation", lambda evaluation_id: (_ for _ in ()).throw(httpx.ReadTimeout("timed out")))
|
||||
with pytest.raises(SystemExit) as exit_:
|
||||
eval_module.eval_crew()
|
||||
assert exit_.value.code == 1
|
||||
assert f"Could not reach AMP while waiting (timed out); the evaluation keeps running at {URL}." in capsys.readouterr().out
|
||||
|
||||
|
||||
def test_a_refusal_mid_poll_names_the_evaluation_and_an_unknown_status_stops_the_wait(project, monkeypatch, capsys):
|
||||
directory, _ = project
|
||||
record_last_run(directory)
|
||||
install(monkeypatch, FakeAMP(statuses=[httpx.Response(404, text="gone")]))
|
||||
with pytest.raises(SystemExit) as exit_:
|
||||
eval_module.eval_crew()
|
||||
assert exit_.value.code == 1 and "AMP answered 404 for evaluation ev-1." in capsys.readouterr().out
|
||||
|
||||
for odd in (httpx.Response(200, json={"id": "ev-1", "status": "cancelled"}), httpx.Response(200, json=[]), httpx.Response(200, text="<html>")):
|
||||
install(monkeypatch, FakeAMP(statuses=[httpx.Response(200, json={"id": "ev-1", "status": "queued"}), odd]))
|
||||
with pytest.raises(SystemExit) as exit_:
|
||||
eval_module.eval_crew()
|
||||
assert exit_.value.code == 1
|
||||
assert f"AMP answered without a known evaluation status" in capsys.readouterr().out
|
||||
|
||||
|
||||
def test_ctrl_c_leaves_the_evaluation_running_and_exits_130(project, monkeypatch, capsys):
|
||||
directory, _ = project
|
||||
record_last_run(directory)
|
||||
amp = install(monkeypatch, FakeAMP())
|
||||
monkeypatch.setattr(amp, "get_evaluation", lambda evaluation_id: (_ for _ in ()).throw(KeyboardInterrupt()))
|
||||
|
||||
with pytest.raises(SystemExit) as exit_:
|
||||
eval_module.eval_crew()
|
||||
|
||||
assert exit_.value.code == 130
|
||||
assert f"Still running at {URL}." in capsys.readouterr().out
|
||||
|
||||
|
||||
def test_dmn_mode_prints_the_url_but_opens_no_browser(project, monkeypatch, capsys):
|
||||
directory, opened = project
|
||||
record_last_run(directory)
|
||||
monkeypatch.setattr(eval_module, "is_dmn_mode_enabled", lambda: True)
|
||||
install(monkeypatch, FakeAMP(statuses=[done()]))
|
||||
|
||||
eval_module.eval_crew()
|
||||
|
||||
assert opened == [] and URL in capsys.readouterr().out
|
||||
|
||||
|
||||
def test_run_skips_the_offer_when_nothing_is_recorded(project, monkeypatch, capsys):
|
||||
monkeypatch.setattr(eval_module.click, "confirm", lambda *args, **kwargs: pytest.fail("no offer with --run"))
|
||||
amp = install(monkeypatch, FakeAMP(statuses=[done()]))
|
||||
|
||||
eval_module.eval_crew(run_id="named-run")
|
||||
|
||||
assert amp.calls[0] == ("create", "named-run")
|
||||
|
||||
|
||||
def test_outside_a_crewai_project_nothing_is_written_and_it_says_so(project, monkeypatch, capsys):
|
||||
directory, _ = project
|
||||
monkeypatch.setattr(eval_module.sys.stdin, "isatty", lambda: True)
|
||||
monkeypatch.setattr(eval_module.click, "confirm", lambda *args, **kwargs: pytest.fail("no offer outside a project"))
|
||||
amp = install(monkeypatch, FakeAMP())
|
||||
|
||||
with pytest.raises(SystemExit) as exit_:
|
||||
eval_module.eval_crew()
|
||||
|
||||
assert exit_.value.code == 1
|
||||
assert "No crewAI project here (no pyproject.toml)" in capsys.readouterr().out
|
||||
assert not (directory / ".env").exists() and amp.calls == []
|
||||
|
||||
|
||||
def test_without_a_traced_run_and_no_terminal_it_explains_and_exits(project, monkeypatch, capsys):
|
||||
(project[0] / "pyproject.toml").write_text("[project]\nname = 'demo'\n")
|
||||
monkeypatch.setattr(eval_module, "is_dmn_mode_enabled", lambda: True)
|
||||
amp = install(monkeypatch, FakeAMP())
|
||||
|
||||
with pytest.raises(SystemExit) as exit_:
|
||||
eval_module.eval_crew()
|
||||
|
||||
assert exit_.value.code == 1
|
||||
out = capsys.readouterr().out
|
||||
assert "No traced run is recorded" in out and "CREWAI_TRACING_ENABLED=true" in out and "crewai run" in out
|
||||
assert amp.calls == []
|
||||
|
||||
|
||||
def test_without_a_traced_run_it_offers_to_turn_tracing_on_and_run_the_crew(project, monkeypatch, capsys):
|
||||
directory, _ = project
|
||||
(directory / "pyproject.toml").write_text("[project]\nname = 'demo'\n")
|
||||
monkeypatch.setattr(eval_module.sys.stdin, "isatty", lambda: True)
|
||||
prompts: list[tuple] = []
|
||||
monkeypatch.setattr(eval_module.click, "confirm", lambda text, **kwargs: prompts.append((text, kwargs)) or True)
|
||||
ran: list[str] = []
|
||||
|
||||
def fake_run_crew() -> None:
|
||||
ran.append("run")
|
||||
assert eval_module.os.environ.get("CREWAI_TRACING_ENABLED") == "true"
|
||||
record_last_run(directory, "fresh-run")
|
||||
|
||||
import crewai_cli.run_crew as run_crew_module
|
||||
|
||||
monkeypatch.setattr(run_crew_module, "run_crew", fake_run_crew)
|
||||
monkeypatch.delenv("CREWAI_TRACING_ENABLED", raising=False)
|
||||
amp = install(monkeypatch, FakeAMP(statuses=[done()]))
|
||||
|
||||
eval_module.eval_crew()
|
||||
|
||||
assert ran == ["run"]
|
||||
text, kwargs = prompts[0]
|
||||
assert "CREWAI_TRACING_ENABLED=true stays in .env" in text and kwargs == {"default": True} # Enter is yes (João's call); the prompt names both effects
|
||||
assert "CREWAI_TRACING_ENABLED=true" in (directory / ".env").read_text()
|
||||
assert amp.calls[0] == ("create", "fresh-run")
|
||||
assert "Tracing is on for this project" in capsys.readouterr().out
|
||||
|
||||
|
||||
def test_declining_the_offer_exits_cleanly_with_the_steps(project, monkeypatch, capsys):
|
||||
(project[0] / "pyproject.toml").write_text("[project]\nname = 'demo'\n")
|
||||
monkeypatch.setattr(eval_module.sys.stdin, "isatty", lambda: True)
|
||||
monkeypatch.setattr(eval_module.click, "confirm", lambda *args, **kwargs: False)
|
||||
amp = install(monkeypatch, FakeAMP())
|
||||
|
||||
with pytest.raises(SystemExit) as exit_:
|
||||
eval_module.eval_crew()
|
||||
|
||||
assert exit_.value.code == 0
|
||||
assert "crewai eval" in capsys.readouterr().out and amp.calls == []
|
||||
|
||||
|
||||
def test_a_run_that_leaves_no_trace_behind_is_explained(project, monkeypatch, capsys):
|
||||
(project[0] / "pyproject.toml").write_text("[project]\nname = 'demo'\n")
|
||||
monkeypatch.setattr(eval_module.sys.stdin, "isatty", lambda: True)
|
||||
monkeypatch.setattr(eval_module.click, "confirm", lambda *args, **kwargs: True)
|
||||
import crewai_cli.run_crew as run_crew_module
|
||||
|
||||
monkeypatch.setattr(run_crew_module, "run_crew", lambda: None)
|
||||
install(monkeypatch, FakeAMP())
|
||||
|
||||
with pytest.raises(SystemExit) as exit_:
|
||||
eval_module.eval_crew()
|
||||
|
||||
assert exit_.value.code == 1
|
||||
out = capsys.readouterr().out
|
||||
assert "no trace was recorded" in out and "older than the version that records the last run" in out
|
||||
|
||||
|
||||
def test_the_cli_command_maps_to_the_implementation(monkeypatch):
|
||||
calls = []
|
||||
monkeypatch.setattr("crewai_cli.cli.eval_crew", lambda **kwargs: calls.append(kwargs))
|
||||
runner = CliRunner()
|
||||
|
||||
assert runner.invoke(eval_command, []).exit_code == 0
|
||||
assert runner.invoke(eval_command, ["--run", EXECUTION_ID]).exit_code == 0
|
||||
assert calls == [{"run_id": None}, {"run_id": EXECUTION_ID}]
|
||||
assert "Evaluate the last traced run" in runner.invoke(eval_command, ["--help"]).output
|
||||
|
||||
|
||||
def test_only_a_missing_login_reads_as_anonymous(monkeypatch, capsys):
|
||||
"""An unreadable credential store is not "anonymous": it is said out loud."""
|
||||
from crewai_cli.authentication.token import AuthError
|
||||
|
||||
monkeypatch.setattr(eval_module, "get_auth_token", lambda: (_ for _ in ()).throw(AuthError("No token found")))
|
||||
assert eval_module.saved_login() is None # not logged in: AMP treats the caller as anonymous
|
||||
|
||||
monkeypatch.setattr(eval_module, "get_auth_token", lambda: "login-token")
|
||||
assert eval_module.saved_login() == "login-token"
|
||||
|
||||
for broken in (OSError(13, "Permission denied"), ValueError("Fernet key must be 32 url-safe base64-encoded bytes.")):
|
||||
monkeypatch.setattr(eval_module, "get_auth_token", lambda error=broken: (_ for _ in ()).throw(error))
|
||||
with pytest.raises(SystemExit) as exit_:
|
||||
eval_module.saved_login()
|
||||
assert exit_.value.code == 1
|
||||
out = capsys.readouterr().out
|
||||
assert "Could not read the saved login" in out and type(broken).__name__ in out and "crewai login" in out
|
||||
|
||||
|
||||
def test_read_last_run_reads_the_record_crewai_writes(tmp_path):
|
||||
assert eval_module.read_last_run(tmp_path) is None
|
||||
(tmp_path / ".crewai").mkdir()
|
||||
(tmp_path / ".crewai" / "last_run.json").write_text("not json")
|
||||
assert eval_module.read_last_run(tmp_path) is None
|
||||
(tmp_path / ".crewai" / "last_run.json").write_text(json.dumps({"tier": "ephemeral"}))
|
||||
assert eval_module.read_last_run(tmp_path) is None
|
||||
record_last_run(tmp_path)
|
||||
record = eval_module.read_last_run(tmp_path)
|
||||
assert record is not None and record["execution_id"] == EXECUTION_ID and record["amp_base_url"] == "https://amp.test"
|
||||
@@ -18,6 +18,33 @@ class TestPlusAPI(unittest.TestCase):
|
||||
self.assertIn("CrewAI-CLI/", self.api.headers["User-Agent"])
|
||||
self.assertTrue(self.api.headers["X-Crewai-Version"])
|
||||
|
||||
@patch("crewai_core.plus_api.PlusAPI._make_request")
|
||||
def test_create_evaluation(self, mock_make_request):
|
||||
mock_response = MagicMock()
|
||||
mock_make_request.return_value = mock_response
|
||||
|
||||
response = self.api.create_evaluation("6f31fe1a-20bd-4bfe-a011-25d6b9341f62")
|
||||
|
||||
mock_make_request.assert_called_once_with(
|
||||
"POST",
|
||||
"/crewai_plus/api/v1/tracing/evaluations",
|
||||
json={"execution_id": "6f31fe1a-20bd-4bfe-a011-25d6b9341f62"},
|
||||
timeout=120.0, # AMP reads the run's spans inside this request
|
||||
)
|
||||
self.assertEqual(response, mock_response)
|
||||
|
||||
@patch("crewai_core.plus_api.PlusAPI._make_request")
|
||||
def test_get_evaluation(self, mock_make_request):
|
||||
mock_response = MagicMock()
|
||||
mock_make_request.return_value = mock_response
|
||||
|
||||
response = self.api.get_evaluation("ev-1")
|
||||
|
||||
mock_make_request.assert_called_once_with(
|
||||
"GET", "/crewai_plus/api/v1/tracing/evaluations/ev-1", timeout=30.0
|
||||
)
|
||||
self.assertEqual(response, mock_response)
|
||||
|
||||
@patch("crewai_core.plus_api.PlusAPI._make_request")
|
||||
def test_login_to_tool_repository(self, mock_make_request):
|
||||
mock_response = MagicMock()
|
||||
|
||||
Reference in New Issue
Block a user