mirror of
https://github.com/crewAIInc/crewAI.git
synced 2026-09-22 10:56:50 +00:00
63 lines
2.2 KiB
Plaintext
63 lines
2.2 KiB
Plaintext
---
|
|
title: Vision Tool
|
|
description: The `VisionTool` is designed to extract text from images.
|
|
icon: eye
|
|
mode: "wide"
|
|
---
|
|
|
|
# `VisionTool`
|
|
|
|
## Description
|
|
|
|
This tool is used to extract text from images. When passed to the agent it will extract the text from the image and then use it to generate a response, report or any other output.
|
|
The URL or the PATH of the image should be passed to the Agent.
|
|
|
|
You can also ask a custom `query` about the image and pick a `complexity_level` that automatically selects the model best suited for the request:
|
|
|
|
| Complexity level | Model |
|
|
| :--------------- | :------------ |
|
|
| `easy` | `gpt-5.6-luna` |
|
|
| `medium` (default) | `gpt-5.6-terra` |
|
|
| `hard` | `gpt-5.6-sol` |
|
|
|
|
When an explicit `llm` or `model` is provided to the tool, it takes precedence over the complexity-based model selection.
|
|
|
|
## Installation
|
|
|
|
Install the crewai_tools package
|
|
|
|
```shell
|
|
pip install 'crewai[tools]'
|
|
```
|
|
|
|
## Usage
|
|
|
|
In order to use the VisionTool, the OpenAI API key should be set in the environment variable `OPENAI_API_KEY`.
|
|
|
|
```python Code
|
|
from crewai_tools import VisionTool
|
|
|
|
vision_tool = VisionTool()
|
|
|
|
@agent
|
|
def researcher(self) -> Agent:
|
|
'''
|
|
This agent uses the VisionTool to extract text from images.
|
|
'''
|
|
return Agent(
|
|
config=self.agents_config["researcher"],
|
|
allow_delegation=False,
|
|
tools=[vision_tool]
|
|
)
|
|
```
|
|
|
|
## Arguments
|
|
|
|
The VisionTool accepts the following arguments:
|
|
|
|
| Argument | Type | Description |
|
|
| :------------------- | :------- | :----------------------------------------------------------------------------------------------------------- |
|
|
| **image_path_url** | `string` | **Mandatory**. The path to the image file (or URL) from which text needs to be extracted. |
|
|
| **query** | `string` | **Optional**. The question or instruction to ask the model about the image. Defaults to `"What's in this image?"`. |
|
|
| **complexity_level** | `string` | **Optional**. The complexity of the request, which selects the model: `easy`, `medium`, or `hard`. Defaults to `medium`. |
|