mirror of
https://github.com/crewAIInc/crewAI.git
synced 2026-09-22 19:06:25 +00:00
feat(cli): crewai eval evaluates the last traced run through AMP (#7649)
* feat(cli): crewai eval evaluates the last traced run through AMP `crewai eval` reads `.crewai/last_run.json` — the record crewAI writes when a traced run's spans reach Wharf — and asks AMP to evaluate that run: POST /crewai_plus/api/v1/tracing/evaluations with the execution id, sending the saved `crewai login` when there is one and nothing otherwise. AMP answers with an evaluation id and a URL; the command prints the URL, opens it, waits for the verdict and prints it (goal gate and the four grades), exit 1 only when the evaluation itself failed. `--run EXECUTION_ID` evaluates another run. With no traced run recorded it offers to turn tracing on for the project (`CREWAI_TRACING_ENABLED=true` in .env, set_key so nothing else in the file moves) and run the crew now with `crewai run`; without a terminal, or declined, it prints the three steps instead. A run that leaves no record behind is explained, never guessed at. AMP's refusals are printed in its own words: a run that needs an account (401 account_required), a refused credential (then `crewai login`), a run AMP does not hold (404), rate limiting (429 with Retry-After), and any other status with AMP's message. The two AMP calls live on the CLI's PlusAPI subclass, so no crewai-core release is needed. Who may evaluate what is AMP's decision, not the command's: an anonymous run once without an account, then it needs one; a run traced while logged in for that organization's members; a deployment execution for members who may see its traces. Co-Authored-By: Claude Fable 5.1 <noreply@anthropic.com> * fix(cli): crewai eval — guard the project, survive AMP blips, name the right subject, read the whole record Review findings, each reproduced before the fix: - The run-it-now offer wrote CREWAI_TRACING_ENABLED into ./.env before checking the directory is a crewAI project, then run_crew() died on a missing pyproject.toml with a traceback. Now: no pyproject.toml → one sentence, exit 1, nothing written. - httpx errors (AMP unreachable, a timeout) surfaced as tracebacks. Now the start says "Could not reach AMP to start the evaluation: …"; while waiting, an unreachable AMP or a 5xx is retried up to POLL_RETRIES consecutive times, then reported with the URL — the evaluation keeps running server-side either way. - The POST now carries a 120 s timeout (AMP reads the run's spans inside it), the poll 30 s. - A 200 whose body has no known status (a non-dict, no status, a status outside queued/running/done/failed) polled forever. Now it stops with the status it saw and the URL. - A 404 without a JSON message read "AMP holds no run <evaluation id>" while polling. _refused takes "run <id>" / "evaluation <id>" and says "AMP answered 404 for <subject>" — AMP's own message still wins. - The record's amp_base_url was ignored; the CLI now evaluates the run at the AMP it was traced to. --run keeps the configured AMP. - The post-run explanation names the third cause: a crewai older than the version that records the last run. Tests for each, plus the previously untested paths: Ctrl-C exits 130, a 2xx without an id, a refusal mid-poll, DMN opens no browser, --run skips the offer. 51 passed in lib/cli/tests (eval + plus_api). Co-Authored-By: Claude Fable 5.1 <noreply@anthropic.com> * fix(cli): crewai eval — the credential goes only to the configured AMP; an explicit yes before running the crew; a done answer needs a well-formed verdict Review findings (CodeRabbit, the PR gate): - The record's amp_base_url was handed to PlusAPI beside the saved login, so a modified .crewai/last_run.json could send the token to any origin. The client is now built from the configured AMP only (CREWAI_PLUS_URL, the saved settings, app.crewai.com); the project's .env is loaded first, as `crewai run` loads it, so the configured AMP is the one the run was traced to. A record naming another address gets a one-line note and no credential. - The offer to turn tracing on and run the crew defaults to no and says the .env change stays; Enter no longer spends a crew run. - A `done` answer whose verdict is missing or malformed (no gate, grades not an object, a grade not an int or null) is a protocol error, exit 1, instead of an INCONCLUSIVE line with exit 0 or an AttributeError. Tests for each: a foreign origin in the record with a saved token, the .env-loaded same-AMP case, seven malformed verdicts, the prompt's text and default. 59 passed (eval + plus_api); ruff and mypy clean. Co-Authored-By: Claude Fable 5.1 <noreply@anthropic.com> * fix(cli): crewai eval — Enter accepts the offer to run the crew (y/n, Y default; the prompt names both effects) João's call (2026-09-20): the confirm is y/n with Y as the default. The prompt still says tracing stays on in .env and that the crew runs now. Co-Authored-By: Claude Fable 5.1 <noreply@anthropic.com> --------- Co-authored-by: Claude Fable 5.1 <noreply@anthropic.com>
This commit is contained in:
@@ -18,6 +18,33 @@ class TestPlusAPI(unittest.TestCase):
|
||||
self.assertIn("CrewAI-CLI/", self.api.headers["User-Agent"])
|
||||
self.assertTrue(self.api.headers["X-Crewai-Version"])
|
||||
|
||||
@patch("crewai_core.plus_api.PlusAPI._make_request")
|
||||
def test_create_evaluation(self, mock_make_request):
|
||||
mock_response = MagicMock()
|
||||
mock_make_request.return_value = mock_response
|
||||
|
||||
response = self.api.create_evaluation("6f31fe1a-20bd-4bfe-a011-25d6b9341f62")
|
||||
|
||||
mock_make_request.assert_called_once_with(
|
||||
"POST",
|
||||
"/crewai_plus/api/v1/tracing/evaluations",
|
||||
json={"execution_id": "6f31fe1a-20bd-4bfe-a011-25d6b9341f62"},
|
||||
timeout=120.0, # AMP reads the run's spans inside this request
|
||||
)
|
||||
self.assertEqual(response, mock_response)
|
||||
|
||||
@patch("crewai_core.plus_api.PlusAPI._make_request")
|
||||
def test_get_evaluation(self, mock_make_request):
|
||||
mock_response = MagicMock()
|
||||
mock_make_request.return_value = mock_response
|
||||
|
||||
response = self.api.get_evaluation("ev-1")
|
||||
|
||||
mock_make_request.assert_called_once_with(
|
||||
"GET", "/crewai_plus/api/v1/tracing/evaluations/ev-1", timeout=30.0
|
||||
)
|
||||
self.assertEqual(response, mock_response)
|
||||
|
||||
@patch("crewai_core.plus_api.PlusAPI._make_request")
|
||||
def test_login_to_tool_repository(self, mock_make_request):
|
||||
mock_response = MagicMock()
|
||||
|
||||
Reference in New Issue
Block a user