mirror of
https://github.com/crewAIInc/crewAI.git
synced 2026-08-10 08:21:54 +00:00
Some checks failed
CodeQL Advanced / Analyze (actions) (push) Has been cancelled
CodeQL Advanced / Analyze (python) (push) Has been cancelled
Check Documentation Broken Links / Check broken links (push) Has been cancelled
Vulnerability Scan / pip-audit (push) Has been cancelled
Build uv cache / build-cache (3.11) (push) Has been cancelled
Build uv cache / build-cache (3.12) (push) Has been cancelled
Build uv cache / build-cache (3.13) (push) Has been cancelled
Build uv cache / build-cache (3.10) (push) Has been cancelled
Nightly Canary Release / Check for new commits (push) Has been cancelled
Nightly Canary Release / Build nightly packages (push) Has been cancelled
Nightly Canary Release / Publish nightly to PyPI (push) Has been cancelled
Mark stale issues and pull requests / stale (push) Has been cancelled
* feat(crewai-tools): add db2 search tool * refactor(crewai-tools): improve db2 search tool implementation * feat(tools): improve DB2VectorSearchTool validation, security, and configurability * docs: add DB2SearchTool documentation * feat: add DB2 search tool * docs: update DB2SearchTool documentation * fix: address CodeRabbit review feedback * fix: validate non-empty filter_by in DB2ToolSchema * chore: trigger CodeRabbit re-review * feat: fortify DB2 tool; fixed JSON response shape, added input guards and config validation * refactor(db2): replace DB2Config with connection_string field * refactor(db2): remove dead _setup_db2 validator and importlib import * refactor(db2): remove dead guard in _connect as _disconnect() is called at the end of every _run, so self.connection is always None when _connect is called next. The 'if not self.connection' guard was dead code. * fix(db2): tighten _validate_identifier regex. Old regex allowed leading digits, multiple periods and dot-only strings (e.g. '.....' passed). * fix(db2): replace __import__ with importlib.import_module in _generate_embedding as keeping openai as a lazy optional import since it is not always required. * perf(db2): cache OpenAI client in _openai_client to avoid re-instantiation as OpenAI(api_key=...) was recreated on every _generate_embedding call. Extract into _get_openai_client() which lazily initialises and caches self._openai_client on first use, reusing it for all subsequent queries. * docs(db2): clarify tool description to mention embedding fallback * docs(db2): update README supported features to clarify embedding behaviour. 'OpenAI embedding fallback' implied it was optional. Replaced with 'Uses a custom embedding function if supplied, otherwise OpenAI embeddings.' * updated both code examples to use the correct import path and public run() method. * feat(crewai-tools): add db2 search tool * refactor(crewai-tools): improve db2 search tool implementation * feat(tools): improve DB2VectorSearchTool validation, security, and configurability * docs: add DB2SearchTool documentation * feat: add DB2 search tool * docs: update DB2SearchTool documentation * fix: address CodeRabbit review feedback * fix: validate non-empty filter_by in DB2ToolSchema * chore: trigger CodeRabbit re-review * feat: fortify DB2 tool; fixed JSON response shape, added input guards and config validation * fix(db2): address ruff and mypy linter errors * style(db2): apply ruff format to db2_search_tool.py * fix(db2-search-tool): address PR review comments - Restore DirectoryReadTool export accidentally removed; add DB2VectorSearchTool and DB2ToolSchema to crewai_tools.tools __init__ and __all__ - Align _ALLOWED_METRICS whitelist with Db2 VECTOR_DISTANCE API: replace DOT_PRODUCT/L2_DISTANCE with EUCLIDEAN_SQUARED/DOT/HAMMING/MANHATTAN - Replace ImportString fields for db2_package/db2_dbi_package with plain Any + lazy importlib.import_module in new _resolve_db2_packages() to avoid Pydantic default-validation gap where strings were never resolved at construction time - Move docs from frozen docs/v1.13.0/ snapshot to docs/edge/en/tools/database-data/ and register in docs/docs.json; update examples to match actual API (connection_string constructor, not DB2Config), correct return format, and align documented distance metrics with the whitelist * fix(db2-search-tool): resolve default and string db2 package imports dynamically * fix(db2-search-tool): export DB2VectorSearchTool and DB2ToolSchema from package-level crewai_tools * docs(db2-search-tool): fix installation command and import path in README --------- Co-authored-by: priyanshu-krishnan1 <priyanshu.krishnan1@ibm.com> Co-authored-by: GeetikaChugh24 <geetika@ibm.com> Co-authored-by: Lorenze Jay <63378463+lorenzejay@users.noreply.github.com> Co-authored-by: Dhruv Chaturvedi <dhruv_insights@Dhruvs-MacBook-Pro.local>
212 lines
5.9 KiB
Plaintext
212 lines
5.9 KiB
Plaintext
---
|
||
title: Db2 Vector Search Tool
|
||
description: Semantic vector search for CrewAI agents using IBM Db2 native VECTOR_DISTANCE capabilities.
|
||
icon: database
|
||
mode: "wide"
|
||
---
|
||
|
||
# `DB2VectorSearchTool`
|
||
|
||
## Description
|
||
|
||
Perform semantic vector similarity searches against IBM Db2 tables using the native `VECTOR_DISTANCE` function.
|
||
Supports configurable distance metrics, OpenAI or custom embeddings, metadata filtering, and result shaping.
|
||
|
||
## Installation
|
||
|
||
```bash
|
||
pip install ibm_db openai
|
||
```
|
||
|
||
Or with uv:
|
||
|
||
```bash
|
||
uv add ibm_db openai
|
||
```
|
||
|
||
## Environment Variables
|
||
|
||
```bash
|
||
OPENAI_API_KEY=your_openai_key # Required when using default OpenAI embeddings
|
||
DB2_CONNECTION_STRING=DATABASE=TESTDB;HOSTNAME=localhost;PORT=50000;PROTOCOL=TCPIP;UID=db2user;PWD=password;
|
||
```
|
||
|
||
## Basic Usage
|
||
|
||
```python
|
||
from crewai import Agent
|
||
from crewai_tools import DB2VectorSearchTool
|
||
|
||
tool = DB2VectorSearchTool(
|
||
connection_string="DATABASE=TESTDB;HOSTNAME=localhost;PORT=50000;PROTOCOL=TCPIP;UID=db2user;PWD=password;",
|
||
table_name="documents",
|
||
vector_column="embedding",
|
||
)
|
||
|
||
agent = Agent(
|
||
role="Research Assistant",
|
||
goal="Find relevant information in documents",
|
||
tools=[tool],
|
||
)
|
||
```
|
||
|
||
## Full Semantic Search Workflow
|
||
|
||
```python
|
||
import os
|
||
from dotenv import load_dotenv
|
||
from crewai import Agent, Task, Crew, Process
|
||
from crewai_tools import DB2VectorSearchTool
|
||
|
||
load_dotenv()
|
||
|
||
db2_tool = DB2VectorSearchTool(
|
||
connection_string=os.getenv("DB2_CONNECTION_STRING"),
|
||
table_name="documents",
|
||
vector_column="embedding",
|
||
return_columns=["content", "category"],
|
||
limit=3,
|
||
distance_metric="COSINE",
|
||
max_distance=0.35,
|
||
)
|
||
|
||
search_agent = Agent(
|
||
role="Senior Semantic Search Agent",
|
||
goal="Find and analyse documents based on semantic search",
|
||
backstory="You are an expert research assistant who can find relevant information using semantic search in a Db2 database.",
|
||
tools=[db2_tool],
|
||
verbose=True,
|
||
)
|
||
|
||
answer_agent = Agent(
|
||
role="Senior Answer Assistant",
|
||
goal="Generate answers based on retrieved context",
|
||
backstory="You are an expert assistant who generates answers from provided context.",
|
||
tools=[db2_tool],
|
||
verbose=True,
|
||
)
|
||
|
||
search_task = Task(
|
||
description="""Search for relevant documents about {query}.
|
||
Include the relevant information found, vector distances, and returned fields.""",
|
||
agent=search_agent,
|
||
)
|
||
|
||
answer_task = Task(
|
||
description="Given the retrieved Db2 context, generate a final answer.",
|
||
agent=answer_agent,
|
||
)
|
||
|
||
crew = Crew(
|
||
agents=[search_agent, answer_agent],
|
||
tasks=[search_task, answer_task],
|
||
process=Process.sequential,
|
||
verbose=True,
|
||
)
|
||
|
||
result = crew.kickoff(inputs={"query": "What is the role of X in the document?"})
|
||
print(result)
|
||
```
|
||
|
||
## Tool Parameters
|
||
|
||
| Parameter | Type | Default | Description |
|
||
|---|---|---|---|
|
||
| `connection_string` | `str` | required | Db2 connection string. Format: `DATABASE=x;HOSTNAME=x;PORT=50000;PROTOCOL=TCPIP;UID=x;PWD=x;` |
|
||
| `table_name` | `str` | `"documents"` | Table to search. Supports `schema.table` notation. |
|
||
| `vector_column` | `str` | `"embedding"` | Column storing the vector embeddings. |
|
||
| `embedding_model` | `str` | `"text-embedding-3-large"` | OpenAI model used when no custom embedding function is provided. |
|
||
| `return_columns` | `list[str]` | `["content"]` | Columns to include in each result. Must contain at least one entry. |
|
||
| `limit` | `int` | `3` | Maximum number of results (1–100). |
|
||
| `distance_metric` | `str` | `"COSINE"` | Db2 distance metric. See supported values below. |
|
||
| `max_distance` | `float \| None` | `None` | Drop results whose distance exceeds this value. |
|
||
| `custom_embedding_fn` | `Callable[[str], list[float]] \| None` | `None` | Custom embedding function. Overrides OpenAI when provided. |
|
||
|
||
## Supported Distance Metrics
|
||
|
||
The following values map directly to the Db2 `VECTOR_DISTANCE` function:
|
||
|
||
- `COSINE`
|
||
- `EUCLIDEAN`
|
||
- `EUCLIDEAN_SQUARED`
|
||
- `DOT`
|
||
- `HAMMING`
|
||
- `MANHATTAN`
|
||
|
||
Reference: [IBM Db2 VECTOR_DISTANCE documentation](https://www.ibm.com/docs/en/db2/12.1.x?topic=functions-vector-distance)
|
||
|
||
## Schema Parameters (per query)
|
||
|
||
| Parameter | Type | Required | Description |
|
||
|---|---|---|---|
|
||
| `query` | `str` | ✅ | The search query. |
|
||
| `filter_by` | `str \| None` | ❌ | Column name for metadata filtering. Must be paired with `filter_value`. |
|
||
| `filter_value` | `Any \| None` | ❌ | Value to filter on. Must be paired with `filter_by`. |
|
||
|
||
## Return Format
|
||
|
||
```json
|
||
{
|
||
"success": true,
|
||
"results": [
|
||
{
|
||
"distance": 0.1401,
|
||
"data": {
|
||
"content": "Document content here",
|
||
"category": "research"
|
||
}
|
||
}
|
||
]
|
||
}
|
||
```
|
||
|
||
On error:
|
||
|
||
```json
|
||
{
|
||
"success": false,
|
||
"error": "Description of what went wrong",
|
||
"error_type": "ExceptionClassName"
|
||
}
|
||
```
|
||
|
||
## Metadata Filtering
|
||
|
||
```python
|
||
result = db2_tool.run(
|
||
query="machine learning",
|
||
filter_by="category",
|
||
filter_value="research",
|
||
)
|
||
```
|
||
|
||
`filter_by` and `filter_value` must always be provided together. Providing only one raises a validation error.
|
||
|
||
## Custom Embeddings
|
||
|
||
Use any embedding model by supplying a `custom_embedding_fn`:
|
||
|
||
```python
|
||
from sentence_transformers import SentenceTransformer
|
||
from crewai_tools import DB2VectorSearchTool
|
||
|
||
model = SentenceTransformer("sentence-transformers/all-MiniLM-L6-v2")
|
||
|
||
def custom_embeddings(text: str) -> list[float]:
|
||
return model.encode(text).tolist()
|
||
|
||
tool = DB2VectorSearchTool(
|
||
connection_string="DATABASE=TESTDB;HOSTNAME=localhost;PORT=50000;PROTOCOL=TCPIP;UID=db2user;PWD=password;",
|
||
table_name="documents",
|
||
custom_embedding_fn=custom_embeddings,
|
||
)
|
||
```
|
||
|
||
When `custom_embedding_fn` is provided, `OPENAI_API_KEY` is not required.
|
||
|
||
## Security Features
|
||
|
||
- SQL identifier validation (table, column names must match `^[A-Za-z][A-Za-z0-9_]*(\.[A-Za-z][A-Za-z0-9_]*)?$`)
|
||
- Parameterised SQL queries — values never interpolated into SQL strings
|
||
- Distance metric whitelist — only valid Db2 metric names accepted
|