Merge branch 'main' into evoisin/fix-prohibited-files-send
This commit is contained in:
@@ -47,6 +47,7 @@ jobs:
|
||||
docker-image-name: 'docker.io/lasuite/conversations-backend:${{ github.sha }}'
|
||||
-
|
||||
name: Build and push
|
||||
if: always()
|
||||
uses: docker/build-push-action@v6
|
||||
with:
|
||||
context: .
|
||||
@@ -86,6 +87,7 @@ jobs:
|
||||
docker-image-name: 'docker.io/lasuite/conversations-frontend:${{ github.sha }}'
|
||||
-
|
||||
name: Build and push
|
||||
if: always()
|
||||
uses: docker/build-push-action@v6
|
||||
with:
|
||||
context: .
|
||||
|
||||
@@ -11,7 +11,9 @@ and this project adheres to
|
||||
### Fixed
|
||||
|
||||
- 🦺(front) Fix send prohibited file types
|
||||
- 🐛(front) fix target blank links in chat #103
|
||||
- 🚑️(posthog) pass str instead of UUID for user PK #134
|
||||
- ⚡️(web-search) keep running when tool call fails #137
|
||||
|
||||
|
||||
## [0.0.7] - 2025-10-28
|
||||
|
||||
@@ -115,6 +115,31 @@ To start all the services, except the frontend container, you can use the follow
|
||||
$ make run-backend
|
||||
```
|
||||
|
||||
**Setup a basic LLM call**
|
||||
|
||||
To be able to use Conversations, you need to configure at least one Large Language Model (LLM) provider.
|
||||
You can do so by setting the appropriate environment variables in the `env.d/development/common` file:
|
||||
|
||||
```ini
|
||||
AI_BASE_URL=http://host.docker.internal:12434/v1/
|
||||
AI_MODEL=gemma3:4b
|
||||
AI_API_KEY=XXX
|
||||
```
|
||||
|
||||
for a local ollama, or by running a local LLM with docker-compose:
|
||||
|
||||
```shellscript
|
||||
$ make create-compose-with-models
|
||||
```
|
||||
|
||||
which will create a `compose.override.yml` file to start a local models `ai/smollm2`
|
||||
which can be changed later by editing the `compose.override.yml` file.
|
||||
|
||||
You will need to call `make run` after changing the `env.d/development/common`
|
||||
or `compose.override.yml` file.
|
||||
|
||||
You can find more information about configuring LLM providers in the [LLM Configuration](docs/llm-configuration.md) documentation.
|
||||
|
||||
**Adding content**
|
||||
|
||||
You can create a basic demo site by running this command:
|
||||
@@ -141,6 +166,18 @@ You first need to create a superuser account:
|
||||
$ make superuser
|
||||
```
|
||||
|
||||
## Documentation 📚
|
||||
|
||||
Additional documentation is available in the `docs/` directory:
|
||||
|
||||
- [LLM Configuration](docs/llm-configuration.md) - Configure Large Language Models and providers
|
||||
- [Attachments](docs/attachments.md) - How to use attachments in conversations
|
||||
- [Tools for Agents](docs/tools.md) - Available tools and how to add new ones
|
||||
- [Environment Variables](docs/env.md) - All available environment variables
|
||||
- [Installation Guide](docs/installation.md) - Deploy on a Kubernetes cluster
|
||||
- [Theming](docs/theming.md) - Customize the application appearance
|
||||
- [Architecture](docs/architecture.md) - Technical architecture overview
|
||||
|
||||
## Licence 📝
|
||||
|
||||
This work is released under the MIT License (see [LICENSE](https://github.com/suitenumerique/conversations/blob/main/LICENSE)).
|
||||
|
||||
@@ -7,8 +7,8 @@ flowchart TD
|
||||
User -- HTTP --> Front("Frontend (NextJS SPA)")
|
||||
Front -- REST API --> Back("Backend (Django)")
|
||||
Front -- OIDC --> Back -- OIDC ---> OIDC("Keycloak / ProConnect")
|
||||
Back -- REST API --> Yserver
|
||||
Back --> DB("Database (PostgreSQL)")
|
||||
Back <--> Celery --> DB
|
||||
Back --> Cache("Cache (Redis)")
|
||||
Back ----> S3("Minio (S3)")
|
||||
Back -- REST API --> LLM("LLM Providers")
|
||||
```
|
||||
|
||||
@@ -0,0 +1,400 @@
|
||||
# Conversation Attachments
|
||||
|
||||
This document describes how conversation attachments work in the Conversations application, including the upload process, security measures, and how documents are processed for use with Large Language Models (LLMs).
|
||||
|
||||
## Table of Contents
|
||||
|
||||
- [Overview](#overview)
|
||||
- [Supported Attachment Types](#supported-attachment-types)
|
||||
- [Architecture & Flow](#architecture--flow)
|
||||
- [High-Level Overview](#high-level-overview)
|
||||
- [Detailed Technical Flow](#detailed-technical-flow)
|
||||
- [Security & Validation](#security--validation)
|
||||
- [MIME Type Validation](#mime-type-validation)
|
||||
- [Malware Detection](#malware-detection)
|
||||
- [Document Processing for LLMs](#document-processing-for-llms)
|
||||
- [Image Attachments](#image-attachments)
|
||||
- [PDF Documents](#pdf-documents)
|
||||
- [Other Document Types](#other-document-types)
|
||||
- [Configuration](#configuration)
|
||||
|
||||
---
|
||||
|
||||
## Overview
|
||||
|
||||
Conversations allows users to attach files to their conversations with the AI assistant. These attachments can be:
|
||||
- **Images** (displayed directly to vision-capable LLMs)
|
||||
- **PDF documents** (sent as document URLs to the LLM)
|
||||
- **Other documents** (converted to text and indexed for semantic search)
|
||||
|
||||
The attachment system uses **S3-compatible object storage** (such as MinIO in development) to store files securely.
|
||||
The backend generates **presigned URLs** that allow the frontend to upload files directly to the storage,
|
||||
without routing the file data through the backend server.
|
||||
|
||||
Note about documents: The system uses a tool called **MarkItDown** to convert various document formats
|
||||
(Word, Excel, PowerPoint, text files, etc.) into Markdown text for processing by LLMs. When at least
|
||||
one non-PDF/image document is attached, the system enables:
|
||||
- a **Retrieval-Augmented Generation (RAG)** search tool to allow the LLM to query relevant sections of the documents.
|
||||
- a **summarization tool** to provide document summaries on user request.
|
||||
⚠️ naive implementation at the moment, needs improvement before being used in production.
|
||||
|
||||
## Supported Attachment Types
|
||||
The following attachment types are supported:
|
||||
- **Images**: `image/png`, `image/jpeg`, `image/gif`, `image/webp`.
|
||||
- **PDF documents**: `application/pdf`
|
||||
- **Other documents**:
|
||||
- Microsoft Word: `application/vnd.openxmlformats-officedocument.wordprocessingml.document`
|
||||
- Microsoft Excel: `application/vnd.openxmlformats-officedocument.spreadsheetml.sheet`
|
||||
- Microsoft PowerPoint: `application/vnd.openxmlformats-officedocument.presentationml.presentation`
|
||||
- Text files: `text/plain`, `text/markdown`, `text/csv`
|
||||
|
||||
**Warning**: The current implementation for PDF expects the LLM to be able to manage them. We need to
|
||||
improve the handling of PDFs in case the LLM cannot process them natively.
|
||||
|
||||
**Todo**:
|
||||
- Add support for more file types and improve document processing workflows.
|
||||
- Allow PDF management via RAG search when the LLM cannot handle them natively.
|
||||
- Allow file type restrictions based on model settings, instead of globally.
|
||||
- Improve the summarization tool to provide better summaries and handle larger documents.
|
||||
- Start file upload right away when the user selects a file, instead of waiting for the user to send the message.
|
||||
|
||||
|
||||
---
|
||||
|
||||
## Architecture & Flow
|
||||
|
||||
### High-Level Overview
|
||||
|
||||
```
|
||||
┌─────────────┐ ┌─────────────┐ ┌─────────────┐ ┌─────────────┐
|
||||
│ Frontend │ │ Backend │ │ S3 Storage │ │ Malware Det.│
|
||||
└──────┬──────┘ └──────┬──────┘ └──────┬──────┘ └──────┬──────┘
|
||||
│ │ │ │
|
||||
│ 1. Create attachment│ │ │
|
||||
├────────────────────>│ │ │
|
||||
│ │ │ │
|
||||
│ 2. Return presigned │ │ │
|
||||
│ URL for upload │ │ │
|
||||
│<────────────────────┤ │ │
|
||||
│ │ │ │
|
||||
│ 3. Upload file │ │ │
|
||||
│ directly to S3 │ │ │
|
||||
├──────────────────────────────────────────>│ │
|
||||
│ │ │ │
|
||||
│ 4. Notify upload │ │ │
|
||||
│ completed │ │ │
|
||||
├────────────────────>│ │ │
|
||||
│ │ │ │
|
||||
│ │ 5. Detect MIME type │ │
|
||||
│ ├────────────────────>│ │
|
||||
│ │ │ │
|
||||
│ │ 6. Scan for malware │ │
|
||||
│ ├──────────────────────────────────────────>│
|
||||
│ │ │ │
|
||||
│ │ 7. Update status │ │
|
||||
│ 8. Return status │<──────────────────────────────────────────┤
|
||||
│<────────────────────┤ │ │
|
||||
│ │ │ │
|
||||
```
|
||||
|
||||
### Detailed Technical Flow
|
||||
|
||||
#### Step 1: Attachment Creation Request
|
||||
|
||||
When a user selects a file to upload, the frontend sends a POST request to create an attachment record:
|
||||
|
||||
**Endpoint**: `POST /api/conversations/{conversation_id}/attachments/`
|
||||
|
||||
**Request payload**:
|
||||
```json
|
||||
{
|
||||
"file_name": "document.pdf",
|
||||
"size": 1048576,
|
||||
"content_type": "application/pdf"
|
||||
}
|
||||
```
|
||||
|
||||
**Backend processing** (`ChatConversationAttachmentViewSet.perform_create`):
|
||||
1. Verifies the user owns the conversation
|
||||
2. Generates a unique UUID for the file
|
||||
3. Creates a storage key: `{conversation_id}/attachments/{uuid}.{extension}`
|
||||
4. Creates a database record with status `PENDING`
|
||||
|
||||
**Response**:
|
||||
```json
|
||||
{
|
||||
"id": "uuid-of-attachment",
|
||||
"key": "conversation-id/attachments/file-id.pdf",
|
||||
"file_name": "document.pdf",
|
||||
"size": 1048576,
|
||||
"upload_state": "pending",
|
||||
"policy": "https://s3.example.com/bucket/...?presigned-params"
|
||||
}
|
||||
```
|
||||
|
||||
The `policy` field contains a **presigned URL** valid for a limited time (configured by `AWS_S3_UPLOAD_POLICY_EXPIRATION`).
|
||||
|
||||
#### Step 2: Direct Upload to S3
|
||||
|
||||
The frontend uses the presigned URL to upload the file directly to S3 storage using a PUT request.
|
||||
|
||||
**Technical details**:
|
||||
- The presigned URL includes authentication parameters
|
||||
- The upload is done with `Content-Type` header matching the file's MIME type
|
||||
- No backend involvement in the data transfer
|
||||
|
||||
#### Step 3: Upload Completion Notification
|
||||
|
||||
After successful upload, the frontend notifies the backend:
|
||||
|
||||
**Endpoint**: `POST /api/conversations/{conversation_id}/attachments/{attachment_id}/upload-ended/`
|
||||
|
||||
**Backend processing** (`ChatConversationAttachmentViewSet.upload_ended`):
|
||||
|
||||
1. **MIME Type Detection** (`chat/views.py`):
|
||||
```python
|
||||
mime_detector = magic.Magic(mime=True)
|
||||
with default_storage.open(attachment.key, "rb") as file:
|
||||
mimetype = mime_detector.from_buffer(file.read(2048))
|
||||
size = file.size
|
||||
```
|
||||
|
||||
Uses `python-magic` to detect the actual MIME type from file content (first 2048 bytes).
|
||||
|
||||
2. **Update attachment status**:
|
||||
- Status: `PENDING` → `ANALYZING`
|
||||
- Store detected MIME type and actual file size
|
||||
|
||||
3. **Trigger Malware Detection**:
|
||||
```python
|
||||
malware_detection.analyse_file(
|
||||
attachment.key,
|
||||
safe_callback="chat.malware_detection.conversation_safe_attachment_callback",
|
||||
unknown_callback="chat.malware_detection.unknown_attachment_callback",
|
||||
unsafe_callback="chat.malware_detection.conversation_unsafe_attachment_callback",
|
||||
conversation_id=conversation_id,
|
||||
)
|
||||
```
|
||||
|
||||
#### Step 4: Malware Detection Callbacks
|
||||
|
||||
The malware detection service (configurable via `MALWARE_DETECTION_BACKEND`) scans the file and calls one of three callbacks:
|
||||
|
||||
**Safe file** (`conversation_safe_attachment_callback`):
|
||||
- Status: `ANALYZING` → `READY`
|
||||
- File is ready for use
|
||||
|
||||
**Unsafe file** (`conversation_unsafe_attachment_callback`):
|
||||
- Status: `ANALYZING` → `SUSPICIOUS`
|
||||
- File is quarantined and not accessible
|
||||
- Security log entry created
|
||||
|
||||
**Unknown status** (`unknown_attachment_callback`):
|
||||
- Handles special cases (e.g., file too large to analyze)
|
||||
- Status: `ANALYZING` → `FILE_TOO_LARGE_TO_ANALYZE`
|
||||
|
||||
---
|
||||
|
||||
## Security & Validation
|
||||
|
||||
For now, the system is not intended to host user-uploaded files for public download.
|
||||
All files are stored in private S3 buckets with presigned URLs for controlled access and only
|
||||
the owner of the conversation/the uploader can access them, so the risk is quite low around bad use of
|
||||
the attachment system.
|
||||
|
||||
Also, the document content is sent to the LLM and does not prevent any prompt injection attacks, which is not
|
||||
an issue specific to the attachment system but to the overall design of LLM-based applications and should be
|
||||
addressed globally. Also for the moment, the system does not have any action tools that could be used to execute
|
||||
malicious code based on document content.
|
||||
|
||||
### Malware Detection
|
||||
|
||||
The malware detection system is **pluggable** and configurable, allowing different backends to be used.
|
||||
By default, a `DummyBackend` is provided that marks all files as safe.
|
||||
|
||||
⚠️ The current implementation does not disallow any file types or status from being used in conversations.
|
||||
This is a potential security risk and should be addressed in future versions.
|
||||
|
||||
---
|
||||
|
||||
## Document Processing for LLMs
|
||||
|
||||
When a user sends a message with attachments, the system processes them differently based on their type:
|
||||
|
||||
### Image Attachments
|
||||
|
||||
**MIME types**: `image/png`, `image/jpeg`, `image/gif`, `image/webp`, etc.
|
||||
|
||||
**Processing flow**:
|
||||
|
||||
1. **URL Conversion**: Local media URLs are converted to presigned S3 URLs before sending to the LLM:
|
||||
```python
|
||||
# From: chat/agents/local_media_url_processors.py
|
||||
content.url = generate_retrieve_policy(key)
|
||||
```
|
||||
|
||||
2. **Sent to LLM**: Images are sent as `ImageUrl` objects in the prompt:
|
||||
```python
|
||||
ImageUrl(
|
||||
url="https://s3.example.com/bucket/key?presigned-params",
|
||||
identifier="file-id.png",
|
||||
)
|
||||
```
|
||||
|
||||
3. **Vision models** can analyze the image content directly.
|
||||
|
||||
4. **Response processing**: After the LLM responds, presigned URLs are converted back to local URLs for storage:
|
||||
```python
|
||||
# Mapping: presigned_url -> /media-key/{conversation_id}/attachments/{file_id}.png
|
||||
```
|
||||
|
||||
### PDF Documents
|
||||
|
||||
**MIME type**: `application/pdf`
|
||||
|
||||
**Processing flow**:
|
||||
|
||||
1. **Direct URL passing**: PDFs are sent as `DocumentUrl` objects :
|
||||
```python
|
||||
DocumentUrl(
|
||||
url="https://s3.example.com/bucket/key?presigned-params",
|
||||
identifier="file-id.pdf",
|
||||
)
|
||||
```
|
||||
|
||||
2. **LLM processing**: Compatible LLMs can:
|
||||
- Extract and read text from PDFs
|
||||
- Understand document structure
|
||||
- Answer questions about the content
|
||||
|
||||
3. **No conversion needed**: PDFs are passed directly without preprocessing.
|
||||
|
||||
### Other Document Types
|
||||
|
||||
**MIME types**: Word documents, Excel spreadsheets, PowerPoint, text files, Markdown, etc.
|
||||
|
||||
**Processing flow**:
|
||||
|
||||
1. **Document parsing**: When a document is uploaded, it's parsed using the `AlbertRagBackend` class.
|
||||
|
||||
2. **Conversion to Markdown**: Documents are converted using **MarkItDown** library or using the "Albert API" for PDFs.
|
||||
|
||||
3. **RAG (Retrieval-Augmented Generation)**:
|
||||
- Converted text is indexed in a vector database
|
||||
- The LLM uses a `document_rag_search` tool to query relevant sections
|
||||
- Only relevant chunks are sent to the LLM to fit context windows
|
||||
|
||||
4. **Summarization tool** if needed.
|
||||
|
||||
### Processing Strategy Decision Tree
|
||||
|
||||
**Decision logic**:
|
||||
- **No documents**: Standard conversation
|
||||
- **Images**: Send as direct (presigned) URLs to the LLM
|
||||
- **Only PDFs**: Send as direct (presigned) URLs to the LLM
|
||||
- **Other documents present**: Enable RAG search tool + convert to Markdown
|
||||
|
||||
---
|
||||
|
||||
## Configuration
|
||||
|
||||
### Environment Variables
|
||||
|
||||
| Variable | Default | Description |
|
||||
|----------------------------------------------|----------------|------------------------------------------------------------|
|
||||
| `ATTACHMENT_MAX_SIZE` | Configurable | Maximum file size in bytes |
|
||||
| `ATTACHMENT_CHECK_UNSAFE_MIME_TYPES_ENABLED` | `True` | Enable/disable MIME type validation |
|
||||
| `AWS_S3_UPLOAD_POLICY_EXPIRATION` | 3600 | Presigned URL expiration (seconds) |
|
||||
| `AWS_S3_RETRIEVE_POLICY_EXPIRATION` | 3600 | Presigned retrieval URL expiration (seconds) |
|
||||
| `AWS_S3_DOMAIN_REPLACE` | None | Alternative S3 domain for presigned URLs (for development) |
|
||||
| `MALWARE_DETECTION_BACKEND` | `DummyBackend` | Malware scanning backend class |
|
||||
| `MALWARE_DETECTION_PARAMETERS` | `{}` | Backend-specific configuration |
|
||||
| `RAG_FILES_ACCEPTED_FORMATS` | See below | List of MIME types accepted for file uploads |
|
||||
|
||||
#### RAG_FILES_ACCEPTED_FORMATS
|
||||
|
||||
This environment variable controls which file types users are allowed to upload as attachments to conversations.
|
||||
|
||||
**Configuration**:
|
||||
- **Type**: List of strings (comma-separated MIME types when using environment variable)
|
||||
- **Default value**: Includes a comprehensive list of document and image formats:
|
||||
- Microsoft Office documents (`.docx`, `.pptx`, `.xlsx`, `.xls`)
|
||||
- Text files (`.txt`, `.csv`)
|
||||
- PDF documents (`.pdf`)
|
||||
- HTML files
|
||||
- Markdown files (`.md`)
|
||||
- Outlook messages (`.msg`)
|
||||
- Images (`.jpeg`, `.png`, `.gif`, `.webp`)
|
||||
|
||||
**Example configuration**:
|
||||
```ini
|
||||
# In environment variable (comma-separated)
|
||||
RAG_FILES_ACCEPTED_FORMATS="application/pdf,text/plain,image/png,image/jpeg"
|
||||
```
|
||||
|
||||
```python
|
||||
# In Django settings (as a Python list)
|
||||
RAG_FILES_ACCEPTED_FORMATS = [
|
||||
"application/pdf",
|
||||
"text/plain",
|
||||
"image/png",
|
||||
"image/jpeg",
|
||||
]
|
||||
```
|
||||
|
||||
**How it's used**:
|
||||
1. **Backend**: The list is exposed via the `/api/v1.0/config/` endpoint as `chat_upload_accept` (MIME types joined with commas)
|
||||
2. **Frontend**: The configuration is used to validate files before upload in the chat interface:
|
||||
- Checks exact MIME type matches
|
||||
- Supports wildcard patterns (e.g., `image/*` for all image types)
|
||||
- Supports file extension patterns (e.g., `.pdf`)
|
||||
3. **User experience**: Files that don't match the accepted formats are rejected with a user-friendly error message
|
||||
|
||||
**Notes**:
|
||||
|
||||
- This setting controls frontend validation only. Backend validation should also be implemented for security.
|
||||
- Future improvements may include per-model file type restrictions.
|
||||
|
||||
### Storage Configuration
|
||||
|
||||
**MinIO (Development)**:
|
||||
```yaml
|
||||
# docker-compose.yml
|
||||
minio:
|
||||
image: minio/minio
|
||||
environment:
|
||||
MINIO_ROOT_USER: minioadmin
|
||||
MINIO_ROOT_PASSWORD: minioadmin
|
||||
command: server /data --console-address ":9001"
|
||||
```
|
||||
|
||||
---
|
||||
|
||||
## Troubleshooting
|
||||
|
||||
### LLM Cannot Access Image/PDF
|
||||
|
||||
**Possible causes**:
|
||||
- Presigned URL has expired
|
||||
- S3 storage is not accessible from the LLM provider
|
||||
- CORS configuration issues
|
||||
|
||||
**Solution**: Check `AWS_S3_RETRIEVE_POLICY_EXPIRATION` and S3 access policies.
|
||||
|
||||
### Document Not Appearing in RAG Search
|
||||
|
||||
**Possible causes**:
|
||||
- Document conversion failed
|
||||
- Vector database indexing failed
|
||||
|
||||
**Check logs**: Look for errors in `DocumentConverter` and RAG backend logs.
|
||||
|
||||
---
|
||||
|
||||
## Related Documentation
|
||||
|
||||
- [Installation Guide](installation.md) - S3 storage setup
|
||||
- [LLM Configuration](llm-configuration.md) - Model capabilities for attachments
|
||||
- [Architecture](architecture.md) - System overview
|
||||
- [Tools](tools.md) - Document search and RAG tools
|
||||
|
||||
+9
-9
@@ -10,7 +10,6 @@ These are the environment variables you can set for the `conversations-backend`
|
||||
|-------------------------------------------------|-----------------------------------------------------------------------------------------------------------------------------------|---------------------------------------------------------|
|
||||
| DJANGO_ALLOWED_HOSTS | allowed hosts | [] |
|
||||
| DJANGO_SECRET_KEY | secret key | |
|
||||
| DJANGO_SERVER_TO_SERVER_API_TOKENS | | [] |
|
||||
| DB_ENGINE | engine to use for database connections | django.db.backends.postgresql_psycopg2 |
|
||||
| DB_NAME | name of the database | conversations |
|
||||
| DB_USER | user to authenticate with | dinum |
|
||||
@@ -24,12 +23,11 @@ These are the environment variables you can set for the `conversations-backend`
|
||||
| AWS_S3_SECRET_ACCESS_KEY | access key for s3 endpoint | |
|
||||
| AWS_S3_REGION_NAME | region name for s3 endpoint | |
|
||||
| AWS_STORAGE_BUCKET_NAME | bucket name for s3 endpoint | conversations-media-storage |
|
||||
| ATTACHMENT_MAX_SIZE | maximum size of document in bytes | 10485760 |
|
||||
| ATTACHMENT_MAX_SIZE | maximum size of document in bytes | 10485760 |
|
||||
| LANGUAGE_CODE | default language | en-us |
|
||||
| API_USERS_LIST_THROTTLE_RATE_SUSTAINED | throttle rate for api | 180/hour |
|
||||
| API_USERS_LIST_THROTTLE_RATE_BURST | throttle rate for api on burst | 30/minute |
|
||||
| SPECTACULAR_SETTINGS_ENABLE_DJANGO_DEPLOY_CHECK | | false |
|
||||
| TRASHBIN_CUTOFF_DAYS | trashbin cutoff | 30 |
|
||||
| DJANGO_EMAIL_BACKEND | email backend library | django.core.mail.backends.smtp.EmailBackend |
|
||||
| DJANGO_EMAIL_BRAND_NAME | brand name for email | |
|
||||
| DJANGO_EMAIL_HOST | host name of email | |
|
||||
@@ -76,12 +74,14 @@ These are the environment variables you can set for the `conversations-backend`
|
||||
| OIDC_USERINFO_FULLNAME_FIELDS | OIDC token claims to create full name | ["first_name", "last_name"] |
|
||||
| OIDC_USERINFO_SHORTNAME_FIELD | OIDC token claims to create shortname | first_name |
|
||||
| ALLOW_LOGOUT_GET_METHOD | Allow get logout method | true |
|
||||
| AI_API_KEY | AI key to be used for AI Base url | |
|
||||
| AI_BASE_URL | OpenAI compatible AI base url | |
|
||||
| AI_MODEL | AI Model to use | |
|
||||
| AI_AGENT_INSTRUCTION | Base instruction for the AI agent | You are a helpful assistant |
|
||||
| Y_PROVIDER_API_KEY | Y provider API key | |
|
||||
| Y_PROVIDER_API_BASE_URL | Y Provider url | |
|
||||
| LLM_CONFIGURATION_FILE_PATH | Path to the LLM configuration JSON file. See [LLM Configuration](llm-configuration.md) for details | <BASE_DIR>/conversations/configuration/llm/default.json |
|
||||
| LLM_DEFAULT_MODEL_HRID | HRID of the model used for conversations | default-model |
|
||||
| LLM_SUMMARIZATION_MODEL_HRID | HRID of the model used for summarization | default-summarization-model |
|
||||
| AI_API_KEY | AI API key to be used for the default provider (used in default LLM configuration, not for production use) | |
|
||||
| AI_BASE_URL | OpenAI compatible AI base URL (used in default LLM configuration, not for production use) | |
|
||||
| AI_MODEL | AI Model name to use (used in default LLM configuration, not for production use) | |
|
||||
| AI_AGENT_INSTRUCTIONS | Base instruction for the AI agent (used in default LLM configuration, not for production use) | You are a helpful assistant. Wrap formulas... |
|
||||
| AI_AGENT_TOOLS | List of enabled tools for the agent (used in default LLM configuration, not for production use) | [] |
|
||||
| CONVERSION_API_ENDPOINT | Conversion API endpoint | convert-markdown |
|
||||
| CONVERSION_API_CONTENT_FIELD | Conversion api content field | content |
|
||||
| CONVERSION_API_TIMEOUT | Conversion api timeout | 30 |
|
||||
|
||||
@@ -9,7 +9,6 @@ backend:
|
||||
DJANGO_CSRF_TRUSTED_ORIGINS: https://conversations.127.0.0.1.nip.io
|
||||
DJANGO_CONFIGURATION: Feature
|
||||
DJANGO_ALLOWED_HOSTS: conversations.127.0.0.1.nip.io
|
||||
DJANGO_SERVER_TO_SERVER_API_TOKENS: secret-api-key
|
||||
DJANGO_SECRET_KEY: AgoodOrAbadKey
|
||||
DJANGO_SETTINGS_MODULE: conversations.settings
|
||||
DJANGO_SUPERUSER_PASSWORD: admin
|
||||
|
||||
@@ -7,7 +7,7 @@ This document is a step-by-step guide that describes how to install Conversation
|
||||
- k8s cluster with an nginx-ingress controller
|
||||
- an OIDC provider (if you don't have one, we provide an example)
|
||||
- a PostgreSQL server (if you don't have one, we provide an example)
|
||||
- a Memcached server (if you don't have one, we provide an example)
|
||||
- a Redis server (if you don't have one, we provide an example)
|
||||
- a S3 bucket (if you don't have one, we provide an example)
|
||||
|
||||
### Test cluster
|
||||
|
||||
@@ -0,0 +1,412 @@
|
||||
# LLM Configuration
|
||||
|
||||
This document describes how to configure Large Language Models (LLMs) in Conversations via the configuration file.
|
||||
|
||||
## Overview
|
||||
|
||||
Conversations uses a JSON configuration file to define LLM models and providers. This approach allows you to:
|
||||
- Configure multiple LLM models from different providers
|
||||
- Switch between models without code changes
|
||||
- Customize model-specific settings like temperature, max tokens, and system prompts
|
||||
- Enable or disable models dynamically
|
||||
|
||||
The overall structure consists of two main sections: `providers` and `models`.
|
||||
Settings for models, provides customization through `settings` and `profile`, which corresponds to the
|
||||
Pydantic AI model settings and profile. While we currently not use those settings extensively,
|
||||
they are available for future use and advanced configurations, please reach us if you face any problem using them.
|
||||
|
||||
## Configuration File Location
|
||||
|
||||
The default LLM configuration file is located at:
|
||||
```
|
||||
src/backend/conversations/configuration/llm/default.json
|
||||
```
|
||||
|
||||
You can override this location by setting the `LLM_CONFIGURATION_FILE_PATH` environment variable, but be careful as
|
||||
this path must be accessible by the backend application _inside the docker image_:
|
||||
``` ini
|
||||
LLM_CONFIGURATION_FILE_PATH=/path/to/your/llm/config.json
|
||||
```
|
||||
|
||||
## Default Behavior
|
||||
|
||||
### Default Configuration
|
||||
|
||||
The default configuration file is useful for local development and running the test, while it can be used
|
||||
in production, we suggest to create a specific one for production and replace the `settings.` values with
|
||||
`environ.` one.
|
||||
|
||||
The default configuration file (`default.json`) includes:
|
||||
|
||||
1. **Two default models**:
|
||||
- `default-model`: The primary conversational model used for chat interactions
|
||||
- `default-summarization-model`: A specialized model for summarizing conversations
|
||||
|
||||
2. **One default provider**:
|
||||
- `default-provider`: An OpenAI-compatible provider that uses environment variables for configuration
|
||||
|
||||
### Environment Variable Integration
|
||||
|
||||
The configuration uses dynamic value resolution with two special prefixes:
|
||||
|
||||
- `settings.VARIABLE_NAME`: Resolves to a Django setting value
|
||||
- `environ.VARIABLE_NAME`: Resolves to an environment variable value
|
||||
|
||||
For example, in the default configuration:
|
||||
```json
|
||||
{
|
||||
"model_name": "settings.AI_MODEL",
|
||||
"system_prompt": "settings.AI_AGENT_INSTRUCTIONS",
|
||||
"tools": "settings.AI_AGENT_TOOLS"
|
||||
}
|
||||
```
|
||||
|
||||
This allows to configure models in tests using the setting override mechanism from Django/Pytest (but might be replaced
|
||||
later with a simple override of the full configuration like it's done in some tests already).
|
||||
|
||||
### Required Environment Variables
|
||||
|
||||
For the default configuration to work, you need to set these environment variables:
|
||||
|
||||
| Variable | Description | Example |
|
||||
|-------------------------------|----------------------------------------|-----------------------------|
|
||||
| `AI_API_KEY` | API key for the default provider | `sk-...` |
|
||||
| `AI_BASE_URL` | Base URL for the OpenAI-compatible API | `https://api.openai.com/v1` |
|
||||
| `AI_MODEL` | Model name to use | `gpt-4o-mini` |
|
||||
|
||||
### Optional Environment Variables
|
||||
|
||||
If you want to customize the agent behavior and tools, you can set these optional environment variables
|
||||
(defaults are provided in the default configuration):
|
||||
|
||||
| Variable | Description | Default |
|
||||
|-------------------------------|----------------------------------------|-------------------|
|
||||
| `AI_AGENT_INSTRUCTIONS` | System prompt for the agent | see `settings.py` |
|
||||
| `AI_AGENT_TOOLS` | List of enabled tools | `[]` |
|
||||
| `SUMMARIZATION_SYSTEM_PROMPT` | Base prompt of the summarization agent | see `settings.py` |
|
||||
|
||||
### Model Selection
|
||||
|
||||
You can configure which models are used for specific tasks via environment variables:
|
||||
|
||||
| Variable | Description | Default |
|
||||
|--------------------------------|------------------------------------------|-------------------------------|
|
||||
| `LLM_DEFAULT_MODEL_HRID` | HRID of the model used for conversations | `default-model` |
|
||||
| `LLM_SUMMARIZATION_MODEL_HRID` | HRID of the model used for summarization | `default-summarization-model` |
|
||||
|
||||
## Configuration Structure
|
||||
|
||||
The configuration file has two main sections:
|
||||
|
||||
### 1. Providers
|
||||
|
||||
Providers define the API endpoints and authentication for LLM services.
|
||||
|
||||
```json
|
||||
{
|
||||
"providers": [
|
||||
{
|
||||
"hrid": "unique-provider-id",
|
||||
"base_url": "https://api.example.com/v1",
|
||||
"api_key": "environ.API_KEY_VAR",
|
||||
"kind": "openai"
|
||||
}
|
||||
]
|
||||
}
|
||||
```
|
||||
|
||||
**Provider Fields:**
|
||||
|
||||
| Field | Type | Required | Description |
|
||||
|------------|--------|----------|---------------------------------------------------------|
|
||||
| `hrid` | string | Yes | Unique identifier for the provider |
|
||||
| `base_url` | string | Yes | API base URL (can use `settings.` or `environ.` prefix) |
|
||||
| `api_key` | string | Yes | API authentication key (use `environ.` here) |
|
||||
| `kind` | string | Yes | Provider type: `openai` or `mistral` |
|
||||
|
||||
### 2. Models
|
||||
|
||||
Models define the LLMs available in your application.
|
||||
|
||||
```json
|
||||
{
|
||||
"models": [
|
||||
{
|
||||
"hrid": "unique-model-id",
|
||||
"model_name": "gpt-4o-mini",
|
||||
"human_readable_name": "GPT-4o Mini",
|
||||
"provider_name": "unique-provider-id",
|
||||
"profile": null,
|
||||
"settings": {},
|
||||
"is_active": true,
|
||||
"icon": null,
|
||||
"system_prompt": "You are a helpful assistant",
|
||||
"tools": []
|
||||
}
|
||||
]
|
||||
}
|
||||
```
|
||||
|
||||
**Model Fields:**
|
||||
|
||||
| Field | Type | Required | Description |
|
||||
|-----------------------|--------------|----------|-----------------------------------------------------------------------------------------------------|
|
||||
| `hrid` | string | Yes | Unique identifier for the model |
|
||||
| `model_name` | string | Yes | Name of the model as recognized by the provider (can use `settings.` or `environ.` prefix) |
|
||||
| `human_readable_name` | string | Yes | Display name shown to users |
|
||||
| `provider_name` | string | No* | Reference to a provider's `hrid` |
|
||||
| `provider` | object | No* | Inline provider definition (alternative to `provider_name`) |
|
||||
| `profile` | object | No | Model-specific capabilities and settings |
|
||||
| `settings` | object | No | Model inference settings (temperature, max_tokens, etc.) |
|
||||
| `is_active` | boolean | Yes | Whether the model is available for use |
|
||||
| `icon` | string/array | No | Base64-encoded icon or array of icon parts |
|
||||
| `system_prompt` | string | Yes | Default system prompt for the model (can use `settings.` or `environ.` prefix) |
|
||||
| `tools` | array | Yes | List of enabled tools for this model (can use `settings.` or `environ.` prefix for the whole array) |
|
||||
| `supports_streaming` | boolean | No | Whether the model supports streaming responses |
|
||||
|
||||
\* Either `provider_name` or `provider` must be set, unless `model_name` is in the format `<provider>:<model>`.
|
||||
|
||||
## Adding New Models
|
||||
|
||||
### Example 1: Adding a New OpenAI Model
|
||||
|
||||
To add a new OpenAI model using the existing default provider:
|
||||
|
||||
```json
|
||||
{
|
||||
"models": [
|
||||
// ...existing models...
|
||||
{
|
||||
"hrid": "gpt-4-turbo",
|
||||
"model_name": "gpt-4-turbo-preview",
|
||||
"human_readable_name": "GPT-4 Turbo",
|
||||
"provider_name": "default-provider",
|
||||
"profile": null,
|
||||
"settings": {
|
||||
"temperature": 0.7,
|
||||
"max_tokens": 4096
|
||||
},
|
||||
"is_active": true,
|
||||
"icon": null,
|
||||
"system_prompt": "You are an expert AI assistant.",
|
||||
"tools": ["web_search_brave_with_document_backend"],
|
||||
"supports_streaming": true
|
||||
}
|
||||
],
|
||||
"providers": [
|
||||
// ...existing providers...
|
||||
]
|
||||
}
|
||||
```
|
||||
|
||||
### Example 2: Adding a Model using Pydantic AI format
|
||||
|
||||
To add a model with a specific provider using the default Pydantic AI format, you don't need to define the provider separately if you use the `model_name` format `<provider>:<model>`.
|
||||
|
||||
1. **Add the model without provider**:
|
||||
|
||||
```json
|
||||
{
|
||||
"models": [
|
||||
{
|
||||
"hrid": "claude-3-opus",
|
||||
"model_name": "anthropic:claude-3-opus-20240229",
|
||||
"human_readable_name": "Claude 3 Opus",
|
||||
"provider_name": null,
|
||||
"profile": null,
|
||||
"settings": {
|
||||
"temperature": 0.7,
|
||||
"max_tokens": 4096
|
||||
},
|
||||
"is_active": true,
|
||||
"icon": null,
|
||||
"system_prompt": "You are Claude, a helpful AI assistant.",
|
||||
"tools": []
|
||||
}
|
||||
],
|
||||
"providers": []
|
||||
}
|
||||
```
|
||||
|
||||
2**Set the environment variable**:
|
||||
|
||||
Pydantic AI expects the API key in an environment variable named `ANTHROPIC_API_KEY` is this example, so set it accordingly:
|
||||
|
||||
```ini
|
||||
ANTHROPIC_API_KEY=your-api-key-here
|
||||
```
|
||||
|
||||
### Example 3: Adding a Mistral Model
|
||||
|
||||
For Mistral AI models using the Etalab platform:
|
||||
|
||||
```json
|
||||
{
|
||||
"models": [
|
||||
{
|
||||
"hrid": "mistral-large",
|
||||
"model_name": "mistral-large-latest",
|
||||
"human_readable_name": "Mistral Large (Etalab)",
|
||||
"provider_name": "mistral-etalab",
|
||||
"profile": null,
|
||||
"settings": {
|
||||
"temperature": 0.5,
|
||||
"max_tokens": 8192
|
||||
},
|
||||
"is_active": true,
|
||||
"icon": null,
|
||||
"system_prompt": "settings.AI_AGENT_INSTRUCTIONS",
|
||||
"tools": ["web_search_brave_with_document_backend"]
|
||||
}
|
||||
],
|
||||
"providers": [
|
||||
{
|
||||
"hrid": "mistral-etalab",
|
||||
"base_url": "https://api.mistral.etalab.gouv.fr/",
|
||||
"api_key": "environ.MISTRAL_ETALAB_API_KEY",
|
||||
"kind": "mistral"
|
||||
}
|
||||
]
|
||||
}
|
||||
```
|
||||
|
||||
### Example 4: Using Inline Provider Definition
|
||||
|
||||
Instead of referencing a provider by name, you can define it inline if you use a unique configuration:
|
||||
|
||||
```json
|
||||
{
|
||||
"models": [
|
||||
{
|
||||
"hrid": "custom-model",
|
||||
"model_name": "custom-model-v1",
|
||||
"human_readable_name": "Custom Model",
|
||||
"provider": {
|
||||
"hrid": "custom-provider-inline",
|
||||
"base_url": "https://custom-api.example.com/v1",
|
||||
"api_key": "environ.CUSTOM_API_KEY",
|
||||
"kind": "openai"
|
||||
},
|
||||
"settings": {},
|
||||
"is_active": true,
|
||||
"icon": null,
|
||||
"system_prompt": "You are a custom assistant.",
|
||||
"tools": []
|
||||
}
|
||||
]
|
||||
}
|
||||
```
|
||||
|
||||
## Advanced Configuration
|
||||
|
||||
### Model Settings
|
||||
|
||||
The `settings` object supports various inference parameters:
|
||||
|
||||
```json
|
||||
{
|
||||
"settings": {
|
||||
"max_tokens": 4096,
|
||||
"temperature": 0.7,
|
||||
"top_p": 0.9,
|
||||
"timeout": 60.0,
|
||||
"parallel_tool_calls": true,
|
||||
"seed": 42,
|
||||
"presence_penalty": 0.0,
|
||||
"frequency_penalty": 0.0,
|
||||
"logit_bias": {},
|
||||
"stop_sequences": [],
|
||||
"extra_headers": {},
|
||||
"extra_body": {}
|
||||
}
|
||||
}
|
||||
```
|
||||
|
||||
### Model Profile
|
||||
|
||||
The `profile` object defines model capabilities:
|
||||
|
||||
```json
|
||||
{
|
||||
"profile": {
|
||||
"supports_tools": true,
|
||||
"supports_json_schema_output": true,
|
||||
"supports_json_object_output": true,
|
||||
"default_structured_output_mode": "json_schema",
|
||||
"thinking_tags": ["<thinking>", "</thinking>"],
|
||||
"ignore_streamed_leading_whitespace": true
|
||||
}
|
||||
}
|
||||
```
|
||||
|
||||
### Available Tools
|
||||
|
||||
Tools can be specified in the `tools` array. Common tools include:
|
||||
- `web_search_brave_with_document_backend`: Web search using Brave API with document processing
|
||||
|
||||
You can also reference the tools list from Django settings:
|
||||
```json
|
||||
{
|
||||
"tools": "settings.AI_AGENT_TOOLS"
|
||||
}
|
||||
```
|
||||
|
||||
### Custom Icons
|
||||
|
||||
Icons can be provided as base64-encoded PNG images. For long strings, you can split them into an array:
|
||||
|
||||
```json
|
||||
{
|
||||
"icon": [
|
||||
"data:image/png;base64,iVBORw0KGgoAAAANSUhEUgAAABwAAAAcCAMAAABF0y+m",
|
||||
"AAAAn1BMVEUALosAKoovTZjw8vb////+9/jlPUniAAziABUAGIWbpsTwq7HhAAAA"
|
||||
]
|
||||
}
|
||||
```
|
||||
|
||||
## Validation
|
||||
|
||||
The configuration is validated when loaded. Common validation errors include:
|
||||
|
||||
- **Provider not found**: A model references a `provider_name` that doesn't exist in the `providers` array
|
||||
- **Missing provider**: Neither `provider_name` nor `provider` is specified, and `model_name` is not in `<provider>:<model>` format
|
||||
- **Environment variable not set**: A value using `environ.` prefix references an undefined environment variable
|
||||
- **Django setting not set**: A value using `settings.` prefix references an undefined Django setting
|
||||
- **Invalid provider kind**: The `kind` field must be either `openai` or `mistral`
|
||||
|
||||
## Testing Your Configuration
|
||||
|
||||
After modifying the configuration file, you can test it by:
|
||||
|
||||
1. **Checking for syntax errors**:
|
||||
```bash
|
||||
python -m json.tool src/backend/conversations/configuration/llm/default.json
|
||||
```
|
||||
|
||||
2. **Starting the application** and checking the logs for validation errors
|
||||
|
||||
3. **Using the Django shell** to load the configuration:
|
||||
```bash
|
||||
./bin/manage shell
|
||||
```
|
||||
```python
|
||||
from django.conf import settings
|
||||
models = settings.LLM_CONFIGURATIONS
|
||||
models.keys() # Should show all model HRIDs
|
||||
```
|
||||
|
||||
## Best Practices
|
||||
|
||||
1. **Use environment variables** for sensitive data like API keys (with `environ.` prefix)
|
||||
2. **Use Django settings** for configurable values that may change between environments (with `settings.` prefix)
|
||||
3. **Keep provider definitions separate** from models to avoid duplication when using multiple models from the same provider
|
||||
4. **Set `is_active: false`** for models you want to keep in the configuration but temporarily disable
|
||||
5. **Use descriptive `hrid` values** that clearly identify the model and provider
|
||||
6. **Document custom configurations** in your deployment documentation
|
||||
7. **Test configuration changes** in a development environment before deploying to production
|
||||
|
||||
## See Also
|
||||
|
||||
- [Environment Variables Documentation](env.md) - For configuring environment variables
|
||||
- [Installation Guide](installation.md) - For deployment instructions
|
||||
|
||||
+20
-20
@@ -14,15 +14,15 @@ Memory is the first bottleneck; CPU matters only when Celery or the Next.js buil
|
||||
|
||||
## 2. Development Environment Memory Requirements
|
||||
|
||||
| Service | Typical use | Rationale / source |
|
||||
|-----------------------|-------------------------------|-----------------------------------------------------------------------------------------|
|
||||
| PostgreSQL | **1 – 2 GB** | `shared_buffers` starting point ≈ 25% RAM ([postgresql.org][1]) |
|
||||
| Keycloak | **≈ 1.3 GB** | 70% of limit for heap + ~300 MB non-heap ([keycloak.org][2]) |
|
||||
| Redis | **≤ 256 MB** | Empty instance ≈ 3 MB; budget 256 MB to allow small datasets ([stackoverflow.com][3]) |
|
||||
| MinIO | **2 GB (dev) / 32 GB (prod)** | Pre-allocates 1–2 GiB; docs recommend 32 GB per host for ≤ 100 Ti storage ([min.io][4]) |
|
||||
| Django API (+ Celery) | **0.8 – 1.5 GB** | Empirical in-house metrics |
|
||||
| Next.js frontend | **0.5 – 1 GB** | Dev build chain |
|
||||
| Nginx | **< 100 MB** | Static reverse-proxy footprint |
|
||||
| Service | Typical use | Rationale / source |
|
||||
|------------------|-------------------------------|-----------------------------------------------------------------------------------------|
|
||||
| PostgreSQL | **1 – 2 GB** | `shared_buffers` starting point ≈ 25% RAM ([postgresql.org][1]) |
|
||||
| Keycloak | **≈ 1.3 GB** | 70% of limit for heap + ~300 MB non-heap ([keycloak.org][2]) |
|
||||
| Redis | **≤ 256 MB** | Empty instance ≈ 3 MB; budget 256 MB to allow small datasets ([stackoverflow.com][3]) |
|
||||
| MinIO | **2 GB (dev) / 32 GB (prod)** | Pre-allocates 1–2 GiB; docs recommend 32 GB per host for ≤ 100 Ti storage ([min.io][4]) |
|
||||
| Django API | **0.8 – 1.5 GB** | Empirical in-house metrics |
|
||||
| Next.js frontend | **0.5 – 1 GB** | Dev build chain |
|
||||
| Nginx | **< 100 MB** | Static reverse-proxy footprint |
|
||||
|
||||
[1]: https://www.postgresql.org/docs/9.1/runtime-config-resource.html "PostgreSQL: Documentation: 9.1: Resource Consumption"
|
||||
[2]: https://www.keycloak.org/high-availability/concepts-memory-and-cpu-sizing "Concepts for sizing CPU and memory resources - Keycloak"
|
||||
@@ -58,7 +58,7 @@ Production deployments differ significantly from development environments. The t
|
||||
| Service | Memory | Notes |
|
||||
|----------------------------------|------------|----------------------------------------|
|
||||
| PostgreSQL | **2 GB** | Core database |
|
||||
| Django API (+ Celery) | **1.5 GB** | Backend services |
|
||||
| Django API | **1.5 GB** | Backend services |
|
||||
| Nginx | **100 MB** | Static files + reverse proxy |
|
||||
| Redis | **256 MB** | Session storage |
|
||||
| **Total (without auth/storage)** | **≈ 4 GB** | External OIDC + object storage assumed |
|
||||
@@ -81,16 +81,16 @@ Production deployments differ significantly from development environments. The t
|
||||
|
||||
## 5. Ports (dev defaults)
|
||||
|
||||
| Port | Service |
|
||||
|-----------|-----------------------|
|
||||
| 3000 | Next.js |
|
||||
| 8071 | Django |
|
||||
| 8080 | Keycloak |
|
||||
| 8083 | Nginx proxy |
|
||||
| 9000/9001 | MinIO |
|
||||
| 15432 | PostgreSQL (main) |
|
||||
| 5433 | PostgreSQL (Keycloak) |
|
||||
| 1081 | Maildev |
|
||||
| Port | Service |
|
||||
|-----------|----------------------------|
|
||||
| 3000 | Next.js |
|
||||
| 8071 | Django |
|
||||
| 8080 | Keycloak |
|
||||
| 8083 | Nginx proxy |
|
||||
| 9000/9001 | MinIO |
|
||||
| 15432 | PostgreSQL (main) |
|
||||
| 5433 | PostgreSQL (Keycloak) |
|
||||
| 1081 | Maildev (currently unused) |
|
||||
|
||||
## 6. Sizing Guidelines
|
||||
|
||||
|
||||
+4
-4
@@ -4,7 +4,7 @@
|
||||
|
||||
To use this feature, simply set the `FRONTEND_CSS_URL` environment variable to the URL of your custom CSS file. For example:
|
||||
|
||||
```javascript
|
||||
```ini
|
||||
FRONTEND_CSS_URL=http://anything/custom-style.css
|
||||
```
|
||||
|
||||
@@ -38,7 +38,7 @@ The footer is configurable from the theme customization file.
|
||||
|
||||
### Settings 🔧
|
||||
|
||||
```shellscript
|
||||
```ini
|
||||
THEME_CUSTOMIZATION_FILE_PATH=<path>
|
||||
```
|
||||
|
||||
@@ -55,10 +55,10 @@ The translations can be partially overridden from the theme customization file.
|
||||
|
||||
### Settings 🔧
|
||||
|
||||
```shellscript
|
||||
```ini
|
||||
THEME_CUSTOMIZATION_FILE_PATH=<path>
|
||||
```
|
||||
|
||||
### Example of JSON
|
||||
|
||||
The json must follow some rules: https://github.com/suitenumerique/conversations/blob/main/src/helm/env.d/dev/configuration/theme/demo.json
|
||||
The json must follow some rules: https://github.com/suitenumerique/conversations/blob/main/src/helm/env.d/dev/configuration/theme/demo.json
|
||||
|
||||
+238
@@ -0,0 +1,238 @@
|
||||
# Tools for the Conversation Agent
|
||||
|
||||
The conversation agent can be extended with various tools that provide additional capabilities such as web search,
|
||||
weather information, and more. We currently only have web search tools, but more tools can be added as needed.
|
||||
This document explains how to configure and use these tools.
|
||||
|
||||
## Overview
|
||||
|
||||
Tools are functions that the LLM can call during a conversation to access external data or perform specific actions.
|
||||
The agent decides when to use these tools based on the user's query and the conversation context.
|
||||
|
||||
## Configuring Tools for a Model
|
||||
|
||||
Tools are configured at the model level in the LLM configuration file.
|
||||
Each model can have its own set of available tools.
|
||||
|
||||
### Configuration File Location
|
||||
|
||||
Read the [LLM Configuration](llm-configuration.md) document to find out where the configuration file is located
|
||||
and how to use it.
|
||||
|
||||
### Example Configuration
|
||||
|
||||
```json
|
||||
{
|
||||
"models": [
|
||||
{
|
||||
"hrid": "default-model",
|
||||
"model_name": "gpt-4",
|
||||
"human_readable_name": "GPT-4 with Tools",
|
||||
"provider_name": "default-provider",
|
||||
"is_active": true,
|
||||
"system_prompt": "You are a helpful assistant.",
|
||||
"tools": [
|
||||
"web_search_brave",
|
||||
"get_current_weather"
|
||||
]
|
||||
}
|
||||
],
|
||||
"providers": [
|
||||
{
|
||||
"hrid": "default-provider",
|
||||
"base_url": "https://api.openai.com/v1",
|
||||
"api_key": "settings.AI_API_KEY",
|
||||
"kind": "openai"
|
||||
}
|
||||
]
|
||||
}
|
||||
```
|
||||
|
||||
The `tools` field accepts either:
|
||||
- A list of tool names: `["tool_name_1", "tool_name_2"]`
|
||||
- A reference to a settings variable: `"settings.AI_AGENT_TOOLS"`
|
||||
|
||||
## Available Tools
|
||||
|
||||
To make a tool available to be in a model's configuration, it must be registered in the tool registry located at
|
||||
`src/backend/chat/tools/__init__.py`.
|
||||
|
||||
This is not dynamic - any changes to the tool registry require a code deployment...
|
||||
We want to add dynamic loading in the future.
|
||||
|
||||
| Tool Name | Description | Documentation |
|
||||
|------------------------------------------|---------------------------------------------------------------|-----------------------------------------------------------------------------|
|
||||
| `get_current_weather` | Fake weather tool for testing purposes | [Details](tools/get_current_weather.md) |
|
||||
| `web_search_tavily` | Web search using Tavily API | [Details](tools/web_search_tavily.md) |
|
||||
| `web_search_brave` | Web search using Brave Search API with optional summarization | [Details](tools/web_search_brave.md) |
|
||||
| `web_search_brave_with_document_backend` | Web search using Brave with RAG-based document processing | [Details](tools/web_search_brave.md#web_search_brave_with_document_backend) |
|
||||
| `web_search_albert_rag` | ⚠️ **Deprecated** - Web search using Albert API with RAG | [Details](tools/web_search_brave.md#deprecated-web_search_albert_rag) |
|
||||
|
||||
## Adding a New Tool
|
||||
|
||||
To add a new tool to the system, follow these steps:
|
||||
|
||||
### 1. Create the Tool Function
|
||||
|
||||
Create a new Python file in `src/backend/chat/tools/` with your tool function. The function should:
|
||||
|
||||
- Have clear type annotations
|
||||
- Include a comprehensive docstring (the LLM uses this to understand when to use the tool)
|
||||
- Accept `RunContext` as the first parameter if it needs access to conversation context
|
||||
- Return appropriate data types
|
||||
|
||||
Example:
|
||||
```python
|
||||
"""My custom tool for the chat agent."""
|
||||
|
||||
from pydantic_ai import RunContext
|
||||
|
||||
def my_custom_tool(ctx: RunContext, param1: str, param2: int) -> dict:
|
||||
"""
|
||||
Brief description of what the tool does.
|
||||
|
||||
The LLM uses this description to decide when to call this tool.
|
||||
|
||||
Args:
|
||||
ctx (RunContext): The run context containing the conversation.
|
||||
param1 (str): Description of parameter 1.
|
||||
param2 (int): Description of parameter 2.
|
||||
|
||||
Returns:
|
||||
dict: Description of the return value.
|
||||
"""
|
||||
# Your implementation here
|
||||
return {"result": "example"}
|
||||
```
|
||||
|
||||
### 2. Register the Tool
|
||||
|
||||
Add your tool to the registry in `src/backend/chat/tools/__init__.py`:
|
||||
|
||||
```python
|
||||
from .my_custom_tool import my_custom_tool
|
||||
|
||||
def get_pydantic_tools_by_name(name: str) -> Tool:
|
||||
"""Get a tool by its name."""
|
||||
tool_dict = {
|
||||
"get_current_weather": Tool(get_current_weather, takes_ctx=False),
|
||||
"web_search_brave": Tool(
|
||||
web_search_brave, takes_ctx=False, prepare=only_if_web_search_enabled
|
||||
),
|
||||
# Add your tool here
|
||||
"my_custom_tool": Tool(
|
||||
my_custom_tool,
|
||||
takes_ctx=True, # Set to True if your tool needs RunContext
|
||||
# prepare=only_if_web_search_enabled # Optional: add conditions
|
||||
),
|
||||
}
|
||||
return tool_dict[name]
|
||||
```
|
||||
|
||||
### 3. Update Imports
|
||||
|
||||
Don't forget to import your tool function at the top of `__init__.py`:
|
||||
|
||||
```python
|
||||
from .my_custom_tool import my_custom_tool
|
||||
```
|
||||
|
||||
### 4. Add to Model Configuration
|
||||
|
||||
Add your tool name to the `tools` list in your LLM configuration file or
|
||||
to the `AI_AGENT_TOOLS` environment variable for local/test purpose.
|
||||
|
||||
## Tool Preparation: Conditional Tool Availability
|
||||
|
||||
Some tools should only be available under certain conditions. The `prepare` parameter in the `Tool` constructor
|
||||
allows you to specify a function that determines whether a tool should be included.
|
||||
|
||||
### The `only_if_web_search_enabled` Prepare Function
|
||||
|
||||
This is a built-in prepare function that checks if web search feature is enabled in the conversation context:
|
||||
|
||||
```python
|
||||
async def only_if_web_search_enabled(ctx, tool_def: ToolDefinition) -> ToolDefinition | None:
|
||||
"""Prepare function to include a tool only if web search is enabled in the context."""
|
||||
return tool_def if ctx.deps.web_search_enabled else None
|
||||
```
|
||||
|
||||
### Usage
|
||||
|
||||
All web search tools use this prepare function:
|
||||
|
||||
```python
|
||||
"web_search_brave": Tool(
|
||||
web_search_brave,
|
||||
takes_ctx=False,
|
||||
prepare=only_if_web_search_enabled
|
||||
),
|
||||
```
|
||||
|
||||
This ensures that web search tools are only available when the user or conversation settings have enabled web search functionality.
|
||||
|
||||
### Creating Custom Prepare Functions
|
||||
|
||||
You can create your own prepare functions for custom conditions:
|
||||
|
||||
```python
|
||||
async def only_if_feature_enabled(ctx, tool_def: ToolDefinition) -> ToolDefinition | None:
|
||||
"""Include tool only if a specific feature is enabled."""
|
||||
return tool_def if ctx.deps.feature_enabled else None
|
||||
```
|
||||
|
||||
## Web Search Enable/Disable
|
||||
|
||||
Web search tools can be toggled on or off based on conversation settings. When web search is disabled:
|
||||
- Web search tools are not included in the agent's available tools
|
||||
- The LLM cannot make web search calls even if it tries
|
||||
- This is enforced by the `only_if_web_search_enabled` prepare function
|
||||
|
||||
The `web_search_enabled` flag is typically set:
|
||||
- Per conversation in the conversation settings
|
||||
- Per user preference
|
||||
- Through admin configuration
|
||||
|
||||
## Best Practices
|
||||
|
||||
1. **Keep tools focused** - Each tool should do one thing well
|
||||
2. **Clear documentation** - The LLM relies on docstrings to understand when to use tools
|
||||
3. **Error handling** - Tools should handle errors gracefully and return meaningful messages
|
||||
4. **Performance** - Be mindful of API rate limits and timeout values
|
||||
5. **Security** - Never log sensitive data (API keys, user data, etc.)
|
||||
6. **Caching** - Use Django's cache framework for expensive operations when appropriate
|
||||
|
||||
## Troubleshooting
|
||||
|
||||
### Tool Not Being Called
|
||||
|
||||
If the LLM isn't calling your tool:
|
||||
- Check that the tool is registered in `get_pydantic_tools_by_name`
|
||||
- Verify the tool is in the model's `tools` configuration
|
||||
- Review the tool's docstring - make it clearer when the tool should be used
|
||||
- Check if any `prepare` function is preventing the tool from being included
|
||||
|
||||
### Tool Errors
|
||||
|
||||
If a tool is throwing errors:
|
||||
- Check the logs for detailed error messages
|
||||
- Verify all required environment variables are set
|
||||
- Ensure the tool's dependencies are installed
|
||||
- Test the tool function independently
|
||||
|
||||
We recommend wrapping external API calls in try/except blocks to handle potential issues gracefully and use
|
||||
the Pydantic AI `ModelRetry` exception to let the LLM manage the errors.
|
||||
|
||||
### Tool Response Issues
|
||||
|
||||
If the LLM isn't using the tool response correctly:
|
||||
- Ensure the return type is clear and well-structured
|
||||
- Consider returning a `ToolReturn` object with metadata
|
||||
- Check if the response format matches what the LLM expects
|
||||
|
||||
## See Also
|
||||
|
||||
- [Web Search Configuration](llm-configuration.md)
|
||||
- [Architecture](architecture.md)
|
||||
- [Environment Variables](env.md)
|
||||
|
||||
@@ -0,0 +1,113 @@
|
||||
# get_current_weather Tool
|
||||
|
||||
## Overview
|
||||
|
||||
The `get_current_weather` tool is a **fake weather tool** designed for testing and demonstration purposes. It does not connect to any real weather API and always returns hardcoded weather data.
|
||||
|
||||
## Purpose
|
||||
|
||||
This tool is useful for:
|
||||
- **Testing** the tool calling functionality of LLMs
|
||||
- **Demonstrating** how tools work without requiring API keys
|
||||
- **Development** and debugging of the agent system
|
||||
- **Example implementation** for creating new tools
|
||||
|
||||
⚠️ **Warning**: This tool should **not** be used in production environments. It always returns fake data regardless of the location or conditions.
|
||||
|
||||
## Configuration
|
||||
|
||||
### Add to Model
|
||||
|
||||
To enable this tool for a model, add it to the `tools` list in your LLM configuration:
|
||||
|
||||
```json
|
||||
{
|
||||
"models": [
|
||||
{
|
||||
"hrid": "my-model",
|
||||
"tools": [
|
||||
"get_current_weather"
|
||||
]
|
||||
}
|
||||
]
|
||||
}
|
||||
```
|
||||
|
||||
Or via environment variable when using local environment settings:
|
||||
```ini
|
||||
AI_AGENT_TOOLS=get_current_weather
|
||||
```
|
||||
|
||||
### No Additional Settings Required
|
||||
|
||||
This tool does not require any API keys, environment variables, or additional configuration.
|
||||
|
||||
## Function Signature
|
||||
|
||||
```python
|
||||
def get_current_weather(location: str, unit: str) -> dict:
|
||||
"""
|
||||
Get the current weather in a given location.
|
||||
|
||||
Args:
|
||||
location (str): The city and state, e.g. San Francisco, CA.
|
||||
unit (str): The unit of temperature, either 'celsius' or 'fahrenheit'.
|
||||
|
||||
Returns:
|
||||
dict: A dictionary containing the location, temperature, and unit.
|
||||
"""
|
||||
```
|
||||
|
||||
## Parameters
|
||||
|
||||
| Parameter | Type | Required | Description |
|
||||
|------------|------|----------|-----------------------------------------------------------------|
|
||||
| `location` | str | Yes | The city and state (e.g., "San Francisco, CA", "Paris, France") |
|
||||
| `unit` | str | Yes | Temperature unit: either "celsius" or "fahrenheit" |
|
||||
|
||||
## Return Value
|
||||
|
||||
Returns a dictionary with the following structure:
|
||||
|
||||
```python
|
||||
{
|
||||
"location": str, # The location that was queried
|
||||
"temperature": int, # Always 22°C or 72°F
|
||||
"unit": str # The unit that was requested
|
||||
}
|
||||
```
|
||||
|
||||
## How the LLM Uses It
|
||||
|
||||
When a user asks about weather, the LLM will:
|
||||
|
||||
1. **Recognize** the weather-related query
|
||||
2. **Extract** the location from the user's message
|
||||
3. **Determine** the appropriate unit (often from context or user preference)
|
||||
4. **Call** the `get_current_weather` tool
|
||||
5. **Receive** the fake weather data
|
||||
6. **Format** a response to the user
|
||||
|
||||
### Example Conversation
|
||||
|
||||
**User**: "What's the weather like in London?"
|
||||
|
||||
**LLM** (internal): *Calls `get_current_weather("London, UK", "celsius")`*
|
||||
|
||||
**Tool Response**:
|
||||
```json
|
||||
{
|
||||
"location": "London, UK",
|
||||
"temperature": 22,
|
||||
"unit": "celsius"
|
||||
}
|
||||
```
|
||||
|
||||
**LLM** (to user): "The current weather in London, UK is 22°C."
|
||||
|
||||
## See Also
|
||||
|
||||
- [Tools Overview](../tools.md)
|
||||
- [Adding a New Tool](../tools.md#adding-a-new-tool)
|
||||
- [Testing Tools](../tools.md#testing-your-tools)
|
||||
|
||||
@@ -0,0 +1,670 @@
|
||||
# Brave Web Search Tools
|
||||
|
||||
## Overview
|
||||
|
||||
The Brave web search tools enable the conversation agent to search the web using the [Brave Search API](https://brave.com/search/api/).
|
||||
Brave Search is a privacy-focused search engine that provides comprehensive web search results.
|
||||
|
||||
This documentation covers three related tools:
|
||||
1. **`web_search_brave`** - Standard web search with optional summarization
|
||||
2. **`web_search_brave_with_document_backend`** - Web search with RAG-based document processing
|
||||
3. **`web_search_albert_rag`** - ⚠️ **Deprecated** - Use `web_search_brave_with_document_backend` instead
|
||||
|
||||
## Table of Contents
|
||||
|
||||
- [Common Configuration](#common-configuration)
|
||||
- [web_search_brave](#web_search_brave)
|
||||
- [web_search_brave_with_document_backend](#web_search_brave_with_document_backend)
|
||||
- [Deprecated: web_search_albert_rag](#deprecated-web_search_albert_rag)
|
||||
- [Comparison](#comparison)
|
||||
- [Best Practices](#best-practices)
|
||||
- [Troubleshooting](#troubleshooting)
|
||||
|
||||
---
|
||||
|
||||
## Common Configuration
|
||||
|
||||
### Prerequisites
|
||||
|
||||
1. **Brave Search API Key**: Sign up at [Brave Search API](https://brave.com/search/api/) to get an API key
|
||||
2. **Environment Variables**: Configure the required settings
|
||||
|
||||
### Common Environment Variables
|
||||
|
||||
All Brave tools share these common settings:
|
||||
|
||||
| Variable | Required | Default | Description |
|
||||
|---------------------|----------|---------|----------------------------------------------------|
|
||||
| `BRAVE_API_KEY` | **Yes** | None | Your Brave Search API key |
|
||||
| `BRAVE_API_TIMEOUT` | No | 5 | API request timeout in seconds |
|
||||
| `BRAVE_MAX_RESULTS` | No | 8 | Maximum number of search results |
|
||||
| `BRAVE_CACHE_TTL` | No | 1800 | Cache time-to-live in seconds (30 minutes) |
|
||||
|
||||
### Search Parameters
|
||||
|
||||
Check on the Brave API documentation for more details on these parameters:
|
||||
|
||||
| Variable | Required | Default | Description |
|
||||
|-------------------------------|----------|------------|---------------------------------------------------|
|
||||
| `BRAVE_SEARCH_COUNTRY` | No | None | Country code for search (e.g., "US", "FR") |
|
||||
| `BRAVE_SEARCH_LANG` | No | None | Language code (e.g., "en", "fr") |
|
||||
| `BRAVE_SEARCH_SAFE_SEARCH` | No | "moderate" | Safe search level: "off", "moderate", or "strict" |
|
||||
| `BRAVE_SEARCH_SPELLCHECK` | No | True | Enable spell checking |
|
||||
| `BRAVE_SEARCH_EXTRA_SNIPPETS` | No | True | Fetch extra snippets from pages |
|
||||
|
||||
|
||||
Note: even if `BRAVE_SEARCH_EXTRA_SNIPPETS` is enabled, the API may not include them if you don't have a plan for this.
|
||||
This is why, in `web_search_brave`, we also fetch the page content ourselves when needed.
|
||||
|
||||
### Configuration Example
|
||||
|
||||
```bash
|
||||
# .env file
|
||||
BRAVE_API_KEY=BSA-your-api-key-here
|
||||
BRAVE_MAX_RESULTS=8
|
||||
BRAVE_MAX_WORKERS=4
|
||||
BRAVE_SEARCH_COUNTRY=US
|
||||
BRAVE_SEARCH_LANG=en
|
||||
BRAVE_SEARCH_SAFE_SEARCH=moderate
|
||||
```
|
||||
|
||||
### Django Settings
|
||||
|
||||
All Brave settings are defined in `src/backend/conversations/brave_settings.py`:
|
||||
|
||||
```python
|
||||
class BraveSettings:
|
||||
"""Brave settings for web_search_brave tool."""
|
||||
|
||||
BRAVE_API_KEY = values.Value(
|
||||
default=None,
|
||||
environ_name="BRAVE_API_KEY",
|
||||
environ_prefix=None,
|
||||
)
|
||||
# ... more settings
|
||||
```
|
||||
|
||||
---
|
||||
|
||||
## web_search_brave
|
||||
|
||||
### Overview
|
||||
|
||||
Standard Brave web search tool with optional LLM-based summarization of page content.
|
||||
|
||||
### Purpose
|
||||
|
||||
- Search the web for up-to-date information
|
||||
- Extract content from web pages
|
||||
- Optionally summarize content using an LLM
|
||||
- Provide structured results with snippets
|
||||
|
||||
### Additional Configuration
|
||||
|
||||
| Variable | Required | Default | Description |
|
||||
|-------------------------------|----------|---------|-------------------------------------------------|
|
||||
| `BRAVE_SUMMARIZATION_ENABLED` | No | False | Enable LLM-based summarization of fetched pages |
|
||||
|
||||
### Function Signature
|
||||
|
||||
```python
|
||||
def web_search_brave(query: str) -> ToolReturn:
|
||||
"""
|
||||
Search the web for up-to-date information
|
||||
|
||||
Args:
|
||||
query (str): The query to search for.
|
||||
|
||||
Returns:
|
||||
ToolReturn: Formatted search results with metadata
|
||||
"""
|
||||
```
|
||||
|
||||
### Return Value
|
||||
|
||||
Returns a `ToolReturn` object with:
|
||||
|
||||
```python
|
||||
ToolReturn(
|
||||
return_value={
|
||||
"0": {
|
||||
"url": "https://example.com/page1",
|
||||
"title": "Example Page Title",
|
||||
"snippets": ["Extracted or summarized content..."]
|
||||
},
|
||||
"1": {
|
||||
"url": "https://example.com/page2",
|
||||
"title": "Another Page",
|
||||
"snippets": ["More content..."]
|
||||
}
|
||||
},
|
||||
metadata={
|
||||
"sources": {
|
||||
"https://example.com/page1",
|
||||
"https://example.com/page2"
|
||||
}
|
||||
}
|
||||
)
|
||||
```
|
||||
|
||||
### How It Works
|
||||
|
||||
1. **Query API**: Sends search query to Brave Search API
|
||||
2. **Receive Results**: Gets list of matching web pages
|
||||
3. **Fetch Content**: For results without extra_snippets:
|
||||
- Fetches the HTML content using `trafilatura`
|
||||
- Extracts the main text content
|
||||
- Caches the extracted content
|
||||
4. **Summarize (Optional)**: If `BRAVE_SUMMARIZATION_ENABLED=True`:
|
||||
- Sends extracted content to summarization agent
|
||||
- Receives concise summary focused on the query
|
||||
5. **Format Results**: Returns structured data with URLs, titles, and snippets
|
||||
|
||||
### Workflow Diagram
|
||||
|
||||
```
|
||||
User Query
|
||||
↓
|
||||
Brave Search API
|
||||
↓
|
||||
Search Results (URLs, titles, descriptions)
|
||||
↓
|
||||
[For each result without snippets]
|
||||
↓
|
||||
Fetch HTML (trafilatura) → Extract Text → Cache
|
||||
↓
|
||||
[If BRAVE_SUMMARIZATION_ENABLED]
|
||||
↓
|
||||
Summarization Agent (LLM)
|
||||
↓
|
||||
Summary Text
|
||||
↓
|
||||
Format & Return
|
||||
```
|
||||
|
||||
### Caching
|
||||
|
||||
Extracted content is cached to avoid repeated fetches:
|
||||
|
||||
```python
|
||||
cache_key = f"web_search_brave:extract:{url}"
|
||||
cache.set(cache_key, document, settings.BRAVE_CACHE_TTL)
|
||||
```
|
||||
|
||||
**Cache Duration**: Controlled by `BRAVE_CACHE_TTL` (default: 30 minutes)
|
||||
|
||||
### Summarization
|
||||
|
||||
When enabled, the tool uses the `SummarizationAgent` to condense page content:
|
||||
|
||||
```python
|
||||
prompt = f"""
|
||||
Based on the following request, summarize the following text in a concise manner,
|
||||
focusing on the key points regarding the user request.
|
||||
The result should be up to 30 lines long.
|
||||
|
||||
<user request>
|
||||
{query}
|
||||
</user request>
|
||||
|
||||
<text to summarize>
|
||||
{text}
|
||||
</text to summarize>
|
||||
"""
|
||||
```
|
||||
|
||||
**Note**: Summarization is costly (additional LLM calls).
|
||||
Use only when necessary, we prefer the document vector search from `web_search_brave_with_document_backend`.
|
||||
|
||||
### Add to Model
|
||||
|
||||
```json
|
||||
{
|
||||
"models": [
|
||||
{
|
||||
"hrid": "my-model",
|
||||
"tools": [
|
||||
"web_search_brave"
|
||||
]
|
||||
}
|
||||
]
|
||||
}
|
||||
```
|
||||
|
||||
### Example Usage
|
||||
|
||||
**User**: "What are the new features in Django 5.0?"
|
||||
|
||||
**Tool Call**: `web_search_brave("Django 5.0 new features")`
|
||||
|
||||
**Tool Response**:
|
||||
```python
|
||||
{
|
||||
"0": {
|
||||
"url": "https://docs.djangoproject.com/en/5.0/releases/5.0/",
|
||||
"title": "Django 5.0 release notes",
|
||||
"snippets": ["Django 5.0 introduces several new features including..."]
|
||||
},
|
||||
# ... more results
|
||||
}
|
||||
```
|
||||
|
||||
### Registration
|
||||
|
||||
```python
|
||||
"web_search_brave": Tool(
|
||||
web_search_brave,
|
||||
takes_ctx=False,
|
||||
prepare=only_if_web_search_enabled
|
||||
)
|
||||
```
|
||||
|
||||
---
|
||||
|
||||
## web_search_brave_with_document_backend
|
||||
|
||||
### Overview
|
||||
|
||||
Advanced Brave web search tool that uses RAG (Retrieval-Augmented Generation)
|
||||
with a document backend for intelligent content processing and retrieval.
|
||||
|
||||
### Purpose
|
||||
|
||||
- Search the web and process results through a RAG system
|
||||
- Store fetched documents in a temporary vector database
|
||||
- Perform semantic search across fetched content
|
||||
- Return the most relevant chunks based on the query
|
||||
|
||||
### Additional Configuration
|
||||
|
||||
| Variable | Required | Default | Description |
|
||||
|-------------------------------------|----------|------------------|----------------------------------------------|
|
||||
| `BRAVE_RAG_WEB_SEARCH_CHUNK_NUMBER` | No | 10 | Number of chunks to retrieve from RAG search |
|
||||
| `RAG_DOCUMENT_SEARCH_BACKEND` | No | AlbertRagBackend | Document backend for RAG processing |
|
||||
|
||||
### Function Signature
|
||||
|
||||
```python
|
||||
def web_search_brave_with_document_backend(ctx: RunContext, query: str) -> ToolReturn:
|
||||
"""
|
||||
Search the web for up-to-date information
|
||||
|
||||
Args:
|
||||
ctx (RunContext): The run context containing the conversation.
|
||||
query (str): The query to search for.
|
||||
|
||||
Returns:
|
||||
ToolReturn: Formatted search results with RAG-enhanced snippets
|
||||
"""
|
||||
```
|
||||
|
||||
### How It Works
|
||||
|
||||
1. **Query API**: Sends search query to Brave Search API
|
||||
2. **Receive Results**: Gets list of matching web pages
|
||||
3. **Create Temporary Collection**: Creates a temporary vector database collection
|
||||
4. **Fetch & Store**: For each result:
|
||||
- Fetches the HTML content
|
||||
- Extracts the main text
|
||||
- Stores in the temporary document backend
|
||||
5. **RAG Search**: Performs semantic search across stored documents
|
||||
6. **Map Results**: Maps RAG chunks back to original search results
|
||||
7. **Format & Return**: Returns structured data with enhanced snippets
|
||||
8. **Cleanup**: Temporary collection is automatically deleted
|
||||
|
||||
### Workflow Diagram
|
||||
|
||||
```
|
||||
User Query
|
||||
↓
|
||||
Brave Search API
|
||||
↓
|
||||
Search Results (URLs)
|
||||
↓
|
||||
Create Temporary Vector Collection
|
||||
↓
|
||||
[For each URL]
|
||||
↓
|
||||
Fetch HTML → Extract Text → Store in Vector DB
|
||||
↓
|
||||
RAG Semantic Search
|
||||
↓
|
||||
Retrieve Most Relevant Chunks
|
||||
↓
|
||||
Map Chunks to Original URLs
|
||||
↓
|
||||
Format & Return
|
||||
↓
|
||||
Delete Temporary Collection
|
||||
```
|
||||
|
||||
### Temporary Collection
|
||||
|
||||
The tool creates a temporary collection with a unique ID:
|
||||
|
||||
```python
|
||||
with document_store_backend.temporary_collection(f"tmp-{uuid.uuid4()}") as document_store:
|
||||
# Fetch and store documents
|
||||
# Perform search
|
||||
# Collection is automatically deleted on exit
|
||||
```
|
||||
|
||||
### RAG Search
|
||||
|
||||
The RAG backend performs semantic search to find the most relevant content:
|
||||
|
||||
```python
|
||||
rag_results = document_store.search(
|
||||
query,
|
||||
results_count=settings.BRAVE_RAG_WEB_SEARCH_CHUNK_NUMBER,
|
||||
)
|
||||
```
|
||||
|
||||
Returns chunks ranked by relevance to the query, not just keyword matching.
|
||||
|
||||
### Token Usage Tracking
|
||||
|
||||
The tool tracks LLM tokens used during RAG processing:
|
||||
|
||||
```python
|
||||
ctx.usage += RunUsage(
|
||||
input_tokens=rag_results.usage.prompt_tokens,
|
||||
output_tokens=rag_results.usage.completion_tokens,
|
||||
)
|
||||
```
|
||||
|
||||
### Document Backend
|
||||
|
||||
The default backend is `AlbertRagBackend`, but you can configure a different one:
|
||||
|
||||
```bash
|
||||
RAG_DOCUMENT_SEARCH_BACKEND=chat.agent_rag.document_rag_backends.custom_backend.CustomBackend
|
||||
```
|
||||
|
||||
### Add to Model
|
||||
|
||||
```json
|
||||
{
|
||||
"models": [
|
||||
{
|
||||
"hrid": "my-model",
|
||||
"tools": [
|
||||
"web_search_brave_with_document_backend"
|
||||
]
|
||||
}
|
||||
]
|
||||
}
|
||||
```
|
||||
|
||||
### Example Usage
|
||||
|
||||
**User**: "Explain the concept of async views in Django"
|
||||
|
||||
**Tool Call**: `web_search_brave_with_document_backend(ctx, "Django async views explained")`
|
||||
|
||||
**Tool Response**:
|
||||
```python
|
||||
{
|
||||
"0": {
|
||||
"url": "https://docs.djangoproject.com/en/stable/topics/async/",
|
||||
"title": "Asynchronous support",
|
||||
"snippets": [
|
||||
"Django has support for writing asynchronous views...",
|
||||
"Async views are declared using Python's async def syntax..."
|
||||
]
|
||||
},
|
||||
# ... more results with relevant chunks
|
||||
}
|
||||
```
|
||||
|
||||
### Registration
|
||||
|
||||
```python
|
||||
"web_search_brave_with_document_backend": Tool(
|
||||
web_search_brave_with_document_backend,
|
||||
takes_ctx=True,
|
||||
prepare=only_if_web_search_enabled,
|
||||
)
|
||||
```
|
||||
|
||||
### Advantages Over Standard web_search_brave
|
||||
|
||||
| Feature | web_search_brave | web_search_brave_with_document_backend |
|
||||
|-------------------|--------------------------------|----------------------------------------|
|
||||
| Content Retrieval | Full page or summary | Semantic chunks |
|
||||
| Relevance | Keyword-based | Semantic similarity |
|
||||
| Token Efficiency | May include irrelevant content | Only relevant chunks |
|
||||
| Processing | Simpler, faster | More intelligent, slower |
|
||||
| Cost | Lower | Higher (RAG processing) |
|
||||
| Best For | General search | Deep research, technical queries |
|
||||
|
||||
---
|
||||
|
||||
## Deprecated: web_search_albert_rag
|
||||
|
||||
### ⚠️ Deprecation Notice
|
||||
|
||||
The `web_search_albert_rag` tool is **deprecated** and should not be used in new implementations.
|
||||
|
||||
**Replacement**: Use `web_search_brave_with_document_backend` instead, which provides:
|
||||
- Better performance
|
||||
- More control over the RAG backend
|
||||
- Temporary collections (no cleanup issues)
|
||||
- Token usage tracking
|
||||
- Parallel processing support
|
||||
|
||||
### Why Deprecated?
|
||||
|
||||
- Limited to Albert API only
|
||||
- No control over document backend
|
||||
- Less flexible than the new approach
|
||||
- Maintenance burden
|
||||
|
||||
### Timeline
|
||||
|
||||
- **Current**: Still functional but not recommended
|
||||
- **Future**: Will be removed in a future version
|
||||
|
||||
---
|
||||
|
||||
## Comparison
|
||||
|
||||
### When to Use Which Tool?
|
||||
|
||||
#### Use `web_search_brave`
|
||||
|
||||
✅ **Best for**:
|
||||
- General web search queries
|
||||
- Quick information retrieval
|
||||
- When speed is important
|
||||
- Lower cost requirements
|
||||
- Simple fact-finding
|
||||
|
||||
❌ **Not ideal for**:
|
||||
- Deep research requiring precise context
|
||||
- Technical documentation queries
|
||||
- When semantic relevance is crucial
|
||||
|
||||
#### Use `web_search_brave_with_document_backend`
|
||||
|
||||
✅ **Best for**:
|
||||
- Complex technical queries
|
||||
- Research requiring precise context
|
||||
- When semantic relevance is important
|
||||
- Questions needing deep understanding
|
||||
- Documentation and how-to queries
|
||||
|
||||
❌ **Not ideal for**:
|
||||
- Simple factual queries
|
||||
- When speed is critical
|
||||
- Budget-constrained scenarios
|
||||
- High-volume usage
|
||||
|
||||
---
|
||||
|
||||
## Best Practices
|
||||
|
||||
### Query Formulation
|
||||
|
||||
Help the LLM formulate effective queries:
|
||||
|
||||
```python
|
||||
# Good queries
|
||||
"Python asyncio tutorial 2024"
|
||||
"Django REST framework authentication"
|
||||
"React hooks best practices"
|
||||
|
||||
# Poor queries
|
||||
"tell me about programming" # Too vague
|
||||
"how do I do the thing with the stuff" # Unclear
|
||||
```
|
||||
|
||||
### Performance Optimization
|
||||
|
||||
#### 1. Optimize Cache
|
||||
|
||||
```bash
|
||||
# Longer cache for stable content
|
||||
BRAVE_CACHE_TTL=3600 # 1 hour
|
||||
|
||||
# Shorter cache for dynamic content
|
||||
BRAVE_CACHE_TTL=300 # 5 minutes
|
||||
```
|
||||
|
||||
#### 2. Control Result Count
|
||||
|
||||
```bash
|
||||
# Fewer results = faster responses
|
||||
BRAVE_MAX_RESULTS=5
|
||||
|
||||
# More results = more comprehensive
|
||||
BRAVE_MAX_RESULTS=10
|
||||
```
|
||||
|
||||
### Summarization Best Practices
|
||||
|
||||
Only enable summarization when needed:
|
||||
|
||||
```bash
|
||||
# Enable for long-form content
|
||||
BRAVE_SUMMARIZATION_ENABLED=True
|
||||
|
||||
# Disable for speed
|
||||
BRAVE_SUMMARIZATION_ENABLED=False
|
||||
```
|
||||
|
||||
**Cost consideration**: Summarization makes additional LLM calls for each result,
|
||||
significantly increasing costs (and execution time).
|
||||
|
||||
### RAG Configuration
|
||||
|
||||
For `web_search_brave_with_document_backend`:
|
||||
|
||||
```bash
|
||||
# More chunks = more context, higher cost
|
||||
BRAVE_RAG_WEB_SEARCH_CHUNK_NUMBER=10
|
||||
|
||||
# Fewer chunks = faster, less context
|
||||
BRAVE_RAG_WEB_SEARCH_CHUNK_NUMBER=5
|
||||
```
|
||||
|
||||
### Search Parameters
|
||||
|
||||
```bash
|
||||
# Localize results
|
||||
BRAVE_SEARCH_COUNTRY=FR
|
||||
BRAVE_SEARCH_LANG=fr
|
||||
|
||||
# Safe search for public deployments
|
||||
BRAVE_SEARCH_SAFE_SEARCH=strict
|
||||
|
||||
# Enable spell check for better results
|
||||
BRAVE_SEARCH_SPELLCHECK=True
|
||||
```
|
||||
|
||||
---
|
||||
|
||||
## Troubleshooting
|
||||
|
||||
### Common Issues
|
||||
|
||||
#### 1. No Results Returned
|
||||
|
||||
**Symptoms**: Empty results or no snippets
|
||||
|
||||
**Causes**:
|
||||
- Query too specific
|
||||
- Content extraction failed
|
||||
- Trafilatura couldn't parse the pages
|
||||
|
||||
**Solutions**:
|
||||
```bash
|
||||
# Enable extra snippets
|
||||
BRAVE_SEARCH_EXTRA_SNIPPETS=True
|
||||
|
||||
# Increase result count
|
||||
BRAVE_MAX_RESULTS=10
|
||||
|
||||
# Check logs for extraction errors
|
||||
```
|
||||
|
||||
#### 2. API Errors
|
||||
|
||||
**Symptoms**: HTTP errors, authentication failures
|
||||
|
||||
**Causes**:
|
||||
- Invalid API key
|
||||
- Rate limit exceeded
|
||||
- API service issues
|
||||
|
||||
**Solutions**:
|
||||
```bash
|
||||
# Verify API key is set
|
||||
echo $BRAVE_API_KEY
|
||||
|
||||
# Check Brave API dashboard for limits
|
||||
# Implement rate limiting in your application
|
||||
```
|
||||
|
||||
#### 3. The tool is not being called
|
||||
**Symptoms**: LLM doesn't use the tool even when appropriate
|
||||
|
||||
**Causes**:
|
||||
- Web search not enabled for the conversation
|
||||
- Tool not in model configuration
|
||||
|
||||
**Solutions**:
|
||||
- Check conversation settings have `web_search_enabled=True`
|
||||
- Verify tool is in the model's `tools` list
|
||||
|
||||
---
|
||||
|
||||
## Security Considerations
|
||||
|
||||
This tool is quite "raw", so be cautious about:
|
||||
- the results returned by the web search
|
||||
- the context size which might be large when not using summarization or RAG if long results are returned
|
||||
- the query content which might include sensitive information
|
||||
- ...
|
||||
|
||||
### Content Validation
|
||||
|
||||
Be aware that fetched content may contain:
|
||||
- Malicious scripts (mitigated by text extraction)
|
||||
- Inappropriate content
|
||||
- Misinformation
|
||||
- Biased information
|
||||
|
||||
The LLM should evaluate sources critically.
|
||||
|
||||
|
||||
---
|
||||
|
||||
## See Also
|
||||
|
||||
- [Tools Overview](../tools.md)
|
||||
- [Tavily Web Search Tool](web_search_tavily.md)
|
||||
- [LLM Configuration](../llm-configuration.md)
|
||||
- [Environment Variables](../env.md)
|
||||
- [Brave Search API Documentation](https://brave.com/search/api/)
|
||||
|
||||
@@ -0,0 +1,370 @@
|
||||
# web_search_tavily Tool
|
||||
|
||||
## Overview
|
||||
|
||||
The `web_search_tavily` tool enables the conversation agent to search the web for up-to-date
|
||||
information using the [Tavily Search API](https://tavily.com/).
|
||||
|
||||
## Purpose
|
||||
|
||||
This tool allows the LLM to:
|
||||
- Access current, real-time information beyond its training data
|
||||
- Answer questions about recent events, news, or developments
|
||||
- Provide factual information with sources
|
||||
- Retrieve specific information from the web
|
||||
|
||||
## Configuration
|
||||
|
||||
### Prerequisites
|
||||
|
||||
1. **Tavily API Key**: Sign up at [Tavily](https://tavily.com/) to get an API key
|
||||
2. **Environment Variables**: Configure the required settings
|
||||
|
||||
### Environment Variables
|
||||
|
||||
| Variable | Required | Default | Description |
|
||||
|----------------------|----------|---------|--------------------------------------------|
|
||||
| `TAVILY_API_KEY` | **Yes** | None | Your Tavily API key |
|
||||
| `TAVILY_MAX_RESULTS` | No | 5 | Maximum number of search results to return |
|
||||
| `TAVILY_API_TIMEOUT` | No | 10 | API request timeout in seconds |
|
||||
|
||||
### Configuration Example
|
||||
|
||||
```bash
|
||||
# .env file
|
||||
TAVILY_API_KEY=tvly-your-api-key-here
|
||||
TAVILY_MAX_RESULTS=5
|
||||
TAVILY_API_TIMEOUT=10
|
||||
```
|
||||
|
||||
### Add to Model
|
||||
|
||||
To enable this tool for a model, add it to the `tools` list in your LLM configuration:
|
||||
|
||||
```json
|
||||
{
|
||||
"models": [
|
||||
{
|
||||
"hrid": "my-model",
|
||||
"tools": [
|
||||
"web_search_tavily"
|
||||
]
|
||||
}
|
||||
]
|
||||
}
|
||||
```
|
||||
|
||||
Or via environment variable when using local environment settings:
|
||||
|
||||
```ini
|
||||
AI_AGENT_TOOLS=web_search_tavily
|
||||
```
|
||||
|
||||
## Function Signature
|
||||
|
||||
```python
|
||||
def web_search_tavily(query: str) -> list[dict]:
|
||||
"""
|
||||
Search the web for up-to-date information
|
||||
|
||||
Args:
|
||||
query (str): The query to search for.
|
||||
|
||||
Returns:
|
||||
list[dict]: A list of search results, each represented as a dictionary.
|
||||
"""
|
||||
```
|
||||
|
||||
## Parameters
|
||||
|
||||
| Parameter | Type | Required | Description |
|
||||
|-----------|------|----------|-------------------------|
|
||||
| `query` | str | Yes | The search query string |
|
||||
|
||||
## Return Value
|
||||
|
||||
Returns a list of dictionaries, each containing:
|
||||
|
||||
```python
|
||||
{
|
||||
"link": str, # URL of the result
|
||||
"title": str, # Title of the page
|
||||
"snippet": str # Content snippet from the page
|
||||
}
|
||||
```
|
||||
|
||||
### Example Return Value
|
||||
|
||||
```python
|
||||
[
|
||||
{
|
||||
"link": "https://example.com/article1",
|
||||
"title": "Introduction to Python",
|
||||
"snippet": "Python is a high-level programming language known for its simplicity..."
|
||||
},
|
||||
{
|
||||
"link": "https://example.com/article2",
|
||||
"title": "Python Best Practices",
|
||||
"snippet": "Follow these best practices to write clean and efficient Python code..."
|
||||
}
|
||||
]
|
||||
```
|
||||
|
||||
## How the LLM Uses It
|
||||
|
||||
When a user asks for current information or specific facts:
|
||||
|
||||
1. **LLM recognizes** the need for external information
|
||||
2. **Formulates** an appropriate search query
|
||||
3. **Calls** `web_search_tavily(query="search terms")`
|
||||
4. **Receives** a list of search results
|
||||
5. **Synthesizes** the information into a response
|
||||
6. **Provides** the answer with source references
|
||||
|
||||
### Example Conversation
|
||||
|
||||
**User**: "What are the latest developments in quantum computing?"
|
||||
|
||||
**LLM** (internal): *Calls `web_search_tavily("latest developments quantum computing 2024")`*
|
||||
|
||||
**Tool Response**:
|
||||
```python
|
||||
[
|
||||
{
|
||||
"link": "https://techcrunch.com/quantum-news",
|
||||
"title": "Major Breakthrough in Quantum Computing",
|
||||
"snippet": "Researchers announced a significant breakthrough..."
|
||||
},
|
||||
# ... more results
|
||||
]
|
||||
```
|
||||
|
||||
**LLM** (to user): "Based on recent sources, there have been several developments in quantum computing.
|
||||
Researchers recently announced a breakthrough in error correction. Additionally, new quantum processors
|
||||
with improved qubit stability have been unveiled..."
|
||||
|
||||
## Implementation Details
|
||||
|
||||
### Source Code
|
||||
|
||||
Located at: `src/backend/chat/tools/web_search_tavily.py`
|
||||
|
||||
```python
|
||||
"""Web search tool using Tavily for the chat agent."""
|
||||
|
||||
from django.conf import settings
|
||||
|
||||
import requests
|
||||
|
||||
|
||||
def web_search_tavily(query: str) -> list[dict]:
|
||||
"""
|
||||
Search the web for up-to-date information
|
||||
|
||||
Args:
|
||||
query (str): The query to search for.
|
||||
|
||||
Returns:
|
||||
list[dict]: A list of search results, each represented as a dictionary.
|
||||
"""
|
||||
url = "https://api.tavily.com/search"
|
||||
data = {
|
||||
"query": query,
|
||||
"api_key": settings.TAVILY_API_KEY,
|
||||
"max_results": settings.TAVILY_MAX_RESULTS,
|
||||
}
|
||||
response = requests.post(url, json=data, timeout=settings.TAVILY_API_TIMEOUT)
|
||||
response.raise_for_status()
|
||||
|
||||
json_response = response.json()
|
||||
|
||||
raw_search_results = json_response.get("results", [])
|
||||
|
||||
return [
|
||||
{
|
||||
"link": result["url"],
|
||||
"title": result.get("title", ""),
|
||||
"snippet": result.get("content"),
|
||||
}
|
||||
for result in raw_search_results
|
||||
]
|
||||
```
|
||||
|
||||
### Registration
|
||||
|
||||
The tool is registered in `src/backend/chat/tools/__init__.py`:
|
||||
|
||||
```python
|
||||
"web_search_tavily": Tool(
|
||||
web_search_tavily,
|
||||
takes_ctx=False,
|
||||
prepare=only_if_web_search_enabled
|
||||
)
|
||||
```
|
||||
|
||||
Note that:
|
||||
- `takes_ctx=False` - This tool doesn't need the conversation context
|
||||
- `prepare=only_if_web_search_enabled` - Only available when web search is enabled
|
||||
|
||||
## Django Settings
|
||||
|
||||
The tool uses these Django settings from `settings.py`:
|
||||
|
||||
```python
|
||||
# Tavily API
|
||||
TAVILY_API_KEY = values.Value(
|
||||
None, # Tavily API key is not set by default
|
||||
environ_name="TAVILY_API_KEY",
|
||||
environ_prefix=None,
|
||||
)
|
||||
TAVILY_MAX_RESULTS = values.PositiveIntegerValue(
|
||||
default=5,
|
||||
environ_name="TAVILY_MAX_RESULTS",
|
||||
environ_prefix=None,
|
||||
)
|
||||
TAVILY_API_TIMEOUT = values.PositiveIntegerValue(
|
||||
default=10, # seconds
|
||||
environ_name="TAVILY_API_TIMEOUT",
|
||||
environ_prefix=None,
|
||||
)
|
||||
```
|
||||
|
||||
## Error Handling
|
||||
|
||||
The tool may raise exceptions in the following cases:
|
||||
|
||||
### Missing API Key
|
||||
```python
|
||||
# If TAVILY_API_KEY is not set
|
||||
AttributeError: 'Settings' object has no attribute 'TAVILY_API_KEY'
|
||||
```
|
||||
|
||||
**Solution**: Set the `TAVILY_API_KEY` environment variable
|
||||
|
||||
### API Errors
|
||||
```python
|
||||
# If the API request fails
|
||||
requests.exceptions.HTTPError: 401 Unauthorized
|
||||
```
|
||||
|
||||
**Possible causes**:
|
||||
- Invalid API key
|
||||
- Exceeded rate limits
|
||||
- API service unavailable
|
||||
|
||||
### Timeout Errors
|
||||
```python
|
||||
# If the request takes too long
|
||||
requests.exceptions.Timeout
|
||||
```
|
||||
|
||||
**Solution**: Increase `TAVILY_API_TIMEOUT` or check network connectivity
|
||||
|
||||
## Best Practices
|
||||
|
||||
### Query Formulation
|
||||
|
||||
The LLM should formulate queries that are:
|
||||
- **Specific and focused** - Better results with targeted queries
|
||||
- **Up-to-date** - Include year or "latest" when relevant
|
||||
- **Clear** - Avoid ambiguous terms
|
||||
- **Concise** - Remove unnecessary words
|
||||
|
||||
Good query examples:
|
||||
- ✅ "quantum computing breakthroughs 2024"
|
||||
- ✅ "latest Python 3.12 features"
|
||||
- ✅ "climate change COP29 outcomes"
|
||||
|
||||
Poor query examples:
|
||||
- ❌ "tell me about stuff happening" (too vague)
|
||||
- ❌ "what is the weather like today in Paris on November 5th 2024 at 3pm" (too specific/long)
|
||||
|
||||
### Rate Limiting
|
||||
|
||||
Be aware of Tavily API rate limits:
|
||||
- Free tier: Limited requests per month
|
||||
- Paid tiers: Higher limits
|
||||
|
||||
Monitor your usage and implement caching if needed.
|
||||
|
||||
### Result Count
|
||||
|
||||
The `TAVILY_MAX_RESULTS` setting controls how many results are returned:
|
||||
- **Lower values (3-5)**: Faster responses, less context for LLM
|
||||
- **Higher values (8-10)**: More comprehensive, but slower and more expensive
|
||||
|
||||
Recommended: **5 results** for most use cases
|
||||
|
||||
## Troubleshooting
|
||||
|
||||
### Tool Not Being Called
|
||||
|
||||
**Symptoms**: LLM doesn't use web search even when appropriate
|
||||
|
||||
**Possible causes**:
|
||||
1. Web search not enabled for the conversation
|
||||
2. Tool not in model configuration
|
||||
3. API key not set
|
||||
|
||||
**Solutions**:
|
||||
1. Check conversation settings have `web_search_enabled=True`
|
||||
2. Verify tool is in the model's `tools` list
|
||||
3. Confirm `TAVILY_API_KEY` is set
|
||||
|
||||
### No Results Returned
|
||||
|
||||
**Symptoms**: Tool returns empty list
|
||||
|
||||
**Possible causes**:
|
||||
1. Query too specific
|
||||
2. No matching results
|
||||
3. API filtering results
|
||||
|
||||
**Solutions**:
|
||||
1. Try broader query terms
|
||||
2. Check Tavily dashboard for query logs
|
||||
3. Review API response in logs
|
||||
|
||||
### Slow Responses
|
||||
|
||||
**Symptoms**: Tool takes a long time to respond
|
||||
|
||||
**Possible causes**:
|
||||
1. Network latency
|
||||
2. Tavily API slow
|
||||
3. Timeout too high
|
||||
|
||||
**Solutions**:
|
||||
1. Check network connectivity
|
||||
2. Monitor Tavily status page
|
||||
3. Adjust `TAVILY_API_TIMEOUT` if needed
|
||||
|
||||
## Security Considerations
|
||||
|
||||
This tool is quite "raw", and was currently only used for test purpose, so be cautious about:
|
||||
- the results returned by the web search
|
||||
- the context size which might be large if many results are returned
|
||||
- the query content which might include sensitive information
|
||||
- ...
|
||||
|
||||
## Performance Optimization
|
||||
|
||||
### Query Optimization
|
||||
|
||||
You may want to help the LLM formulate better queries by including something like this in the system prompt:
|
||||
|
||||
```
|
||||
When using web search:
|
||||
- Use specific, focused queries
|
||||
- Include relevant time periods if needed
|
||||
- Avoid unnecessary words
|
||||
- Combine related terms
|
||||
```
|
||||
|
||||
## See Also
|
||||
|
||||
- [Tools Overview](../tools.md)
|
||||
- [Brave Web Search Tool](web_search_brave.md)
|
||||
- [Web Search Configuration](../llm-configuration.md)
|
||||
- [Environment Variables](../env.md)
|
||||
|
||||
@@ -1,6 +1,5 @@
|
||||
# For the CI job test-e2e
|
||||
BURST_THROTTLE_RATES="200/minute"
|
||||
DJANGO_SERVER_TO_SERVER_API_TOKENS=test-e2e
|
||||
SUSTAINED_THROTTLE_RATES="200/hour"
|
||||
|
||||
# Features
|
||||
|
||||
@@ -8,6 +8,7 @@ from urllib.parse import urljoin
|
||||
|
||||
from django.conf import settings
|
||||
|
||||
import httpx
|
||||
import requests
|
||||
|
||||
from chat.agent_rag.albert_api_constants import Searches
|
||||
@@ -65,6 +66,27 @@ class AlbertRagBackend(BaseRagBackend): # pylint: disable=too-many-instance-att
|
||||
self.collection_id = str(response.json()["id"])
|
||||
return self.collection_id
|
||||
|
||||
async def acreate_collection(self, name: str, description: Optional[str] = None) -> str:
|
||||
"""
|
||||
Create a temporary collection for the search operation.
|
||||
This method should handle the logic to create or retrieve an existing collection.
|
||||
"""
|
||||
async with httpx.AsyncClient(timeout=settings.ALBERT_API_TIMEOUT) as client:
|
||||
response = await client.post(
|
||||
self._collections_endpoint,
|
||||
headers=self._headers,
|
||||
json={
|
||||
"name": name,
|
||||
"description": description or self._default_collection_description,
|
||||
"visibility": "private",
|
||||
},
|
||||
timeout=settings.ALBERT_API_TIMEOUT,
|
||||
)
|
||||
response.raise_for_status()
|
||||
|
||||
self.collection_id = str(response.json()["id"])
|
||||
return self.collection_id
|
||||
|
||||
def delete_collection(self) -> None:
|
||||
"""
|
||||
Delete the current collection
|
||||
@@ -76,6 +98,18 @@ class AlbertRagBackend(BaseRagBackend): # pylint: disable=too-many-instance-att
|
||||
)
|
||||
response.raise_for_status()
|
||||
|
||||
async def adelete_collection(self) -> None:
|
||||
"""
|
||||
Asynchronously delete the current collection
|
||||
"""
|
||||
async with httpx.AsyncClient(timeout=settings.ALBERT_API_TIMEOUT) as client:
|
||||
response = await client.delete(
|
||||
urljoin(f"{self._collections_endpoint}/", self.collection_id),
|
||||
headers=self._headers,
|
||||
timeout=settings.ALBERT_API_TIMEOUT,
|
||||
)
|
||||
response.raise_for_status()
|
||||
|
||||
def parse_pdf_document(self, name: str, content_type: str, content: BytesIO) -> str:
|
||||
"""
|
||||
Parse the PDF document content and return the text content.
|
||||
@@ -150,6 +184,31 @@ class AlbertRagBackend(BaseRagBackend): # pylint: disable=too-many-instance-att
|
||||
logger.debug(response.json())
|
||||
response.raise_for_status()
|
||||
|
||||
async def astore_document(self, name: str, content: str) -> None:
|
||||
"""
|
||||
Store the document content in the Albert collection.
|
||||
This method should handle the logic to send the document content to the Albert API.
|
||||
|
||||
Args:
|
||||
name (str): The name of the document.
|
||||
content (str): The content of the document in Markdown format.
|
||||
"""
|
||||
async with httpx.AsyncClient(timeout=settings.ALBERT_API_TIMEOUT) as client:
|
||||
response = await client.post(
|
||||
urljoin(self._base_url, self._documents_endpoint),
|
||||
headers=self._headers,
|
||||
files={
|
||||
"file": (f"{name}.md", BytesIO(content.encode("utf-8")), "text/markdown"),
|
||||
},
|
||||
data={
|
||||
"collection": int(self.collection_id),
|
||||
"metadata": json.dumps({"document_name": name}), # undocumented API
|
||||
},
|
||||
timeout=settings.ALBERT_API_TIMEOUT,
|
||||
)
|
||||
logger.debug(response.json())
|
||||
response.raise_for_status()
|
||||
|
||||
def search(self, query, results_count: int = 4) -> RAGWebResults:
|
||||
"""
|
||||
Perform a search using the Albert API based on the provided query.
|
||||
@@ -190,3 +249,48 @@ class AlbertRagBackend(BaseRagBackend): # pylint: disable=too-many-instance-att
|
||||
completion_tokens=searches.usage.completion_tokens,
|
||||
),
|
||||
)
|
||||
|
||||
async def asearch(self, query, results_count: int = 4) -> RAGWebResults:
|
||||
"""
|
||||
Perform an asynchronous search using the Albert API based on the provided query.
|
||||
|
||||
Args:
|
||||
query (str): The search query.
|
||||
results_count (int): The number of results to return.
|
||||
|
||||
Returns:
|
||||
RAGWebResults: The search results.
|
||||
"""
|
||||
async with httpx.AsyncClient(timeout=settings.ALBERT_API_TIMEOUT) as client:
|
||||
response = await client.post(
|
||||
urljoin(self._base_url, self._search_endpoint),
|
||||
headers=self._headers,
|
||||
json={
|
||||
"collections": [int(self.collection_id)],
|
||||
"prompt": query,
|
||||
"score_threshold": 0.6,
|
||||
"k": results_count, # Number of chunks to return from the search
|
||||
},
|
||||
timeout=settings.ALBERT_API_TIMEOUT,
|
||||
)
|
||||
|
||||
logger.debug("Search response: %s %s", response.text, response.status_code)
|
||||
|
||||
response.raise_for_status()
|
||||
|
||||
searches = Searches(**response.json())
|
||||
|
||||
return RAGWebResults(
|
||||
data=[
|
||||
RAGWebResult(
|
||||
url=result.chunk.metadata["document_name"],
|
||||
content=result.chunk.content,
|
||||
score=result.score,
|
||||
)
|
||||
for result in searches.data
|
||||
],
|
||||
usage=RAGWebUsage(
|
||||
prompt_tokens=searches.usage.prompt_tokens,
|
||||
completion_tokens=searches.usage.completion_tokens,
|
||||
),
|
||||
)
|
||||
|
||||
@@ -1,10 +1,12 @@
|
||||
"""Implementation of the Albert API for RAG document search."""
|
||||
|
||||
import logging
|
||||
from contextlib import contextmanager
|
||||
from contextlib import asynccontextmanager, contextmanager
|
||||
from io import BytesIO
|
||||
from typing import Optional
|
||||
|
||||
from asgiref.sync import sync_to_async
|
||||
|
||||
from chat.agent_rag.constants import RAGWebResults
|
||||
|
||||
logger = logging.getLogger(__name__)
|
||||
@@ -25,6 +27,13 @@ class BaseRagBackend:
|
||||
"""
|
||||
raise NotImplementedError("Must be implemented in subclass.")
|
||||
|
||||
async def acreate_collection(self, name: str, description: Optional[str] = None) -> str:
|
||||
"""
|
||||
Create a temporary collection for the search operation.
|
||||
This method should handle the logic to create or retrieve an existing collection.
|
||||
"""
|
||||
return await sync_to_async(self.create_collection)(name=name, description=description)
|
||||
|
||||
def parse_document(self, name: str, content_type: str, content: BytesIO):
|
||||
"""
|
||||
Parse the document and prepare it for the search operation.
|
||||
@@ -43,8 +52,8 @@ class BaseRagBackend:
|
||||
|
||||
def store_document(self, name: str, content: str) -> None:
|
||||
"""
|
||||
Store the document content in the Albert collection.
|
||||
This method should handle the logic to send the document content to the Albert API.
|
||||
Store the document content in the collection.
|
||||
This method should handle the logic to send the document content to the API.
|
||||
|
||||
Args:
|
||||
name (str): The name of the document.
|
||||
@@ -52,6 +61,17 @@ class BaseRagBackend:
|
||||
"""
|
||||
raise NotImplementedError("Must be implemented in subclass.")
|
||||
|
||||
async def astore_document(self, name: str, content: str) -> None:
|
||||
"""
|
||||
Store the document content in the collection.
|
||||
This method should handle the logic to send the document content to the API.
|
||||
|
||||
Args:
|
||||
name (str): The name of the document.
|
||||
content (str): The content of the document in Markdown format.
|
||||
"""
|
||||
return await sync_to_async(self.store_document)(name=name, content=content)
|
||||
|
||||
def parse_and_store_document(self, name: str, content_type: str, content: BytesIO) -> str:
|
||||
"""
|
||||
Parse the document and store it in the Albert collection.
|
||||
@@ -75,12 +95,25 @@ class BaseRagBackend:
|
||||
"""
|
||||
raise NotImplementedError("Must be implemented in subclass.")
|
||||
|
||||
async def adelete_collection(self) -> None:
|
||||
"""
|
||||
Delete the collection.
|
||||
This method should handle the logic to delete the collection from the backend.
|
||||
"""
|
||||
return await sync_to_async(self.delete_collection)()
|
||||
|
||||
def search(self, query, results_count: int = 4) -> RAGWebResults:
|
||||
"""
|
||||
Search the collection for the given query.
|
||||
"""
|
||||
raise NotImplementedError("Must be implemented in subclass.")
|
||||
|
||||
async def asearch(self, query, results_count: int = 4) -> RAGWebResults:
|
||||
"""
|
||||
Search the collection for the given query.
|
||||
"""
|
||||
return await sync_to_async(self.search)(query=query, results_count=results_count)
|
||||
|
||||
@classmethod
|
||||
@contextmanager
|
||||
def temporary_collection(cls, name: str, description: Optional[str] = None):
|
||||
@@ -92,3 +125,15 @@ class BaseRagBackend:
|
||||
yield backend
|
||||
finally:
|
||||
backend.delete_collection()
|
||||
|
||||
@classmethod
|
||||
@asynccontextmanager
|
||||
async def temporary_collection_async(cls, name: str, description: Optional[str] = None):
|
||||
"""Context manager for RAG backend with temporary collections."""
|
||||
backend = cls()
|
||||
|
||||
await backend.acreate_collection(name=name, description=description)
|
||||
try:
|
||||
yield backend
|
||||
finally:
|
||||
await backend.adelete_collection()
|
||||
|
||||
@@ -0,0 +1,154 @@
|
||||
"""Tests for chat tool utilities."""
|
||||
|
||||
import inspect
|
||||
from typing import get_type_hints
|
||||
|
||||
import pytest
|
||||
from pydantic_ai import ModelRetry, RunContext
|
||||
|
||||
from chat.tools.exceptions import ModelCannotRetry
|
||||
from chat.tools.utils import last_model_retry_soft_fail
|
||||
|
||||
|
||||
def test_last_model_retry_soft_fail_preserves_function_metadata():
|
||||
"""Test that the decorator preserves function metadata for schema generation."""
|
||||
|
||||
@last_model_retry_soft_fail
|
||||
async def example_tool(ctx: RunContext, query: str, limit: int = 10) -> str: # pylint: disable=unused-argument
|
||||
"""
|
||||
Example tool function.
|
||||
|
||||
Args:
|
||||
ctx: The run context.
|
||||
query: The search query.
|
||||
limit: Maximum number of results.
|
||||
|
||||
Returns:
|
||||
The search results.
|
||||
"""
|
||||
return f"Results for {query} (limit: {limit})"
|
||||
|
||||
# Check that function name is preserved
|
||||
assert example_tool.__name__ == "example_tool"
|
||||
|
||||
# Check that docstring is preserved
|
||||
assert example_tool.__doc__ is not None
|
||||
assert "Example tool function" in example_tool.__doc__
|
||||
|
||||
# Check that signature is preserved
|
||||
sig = inspect.signature(example_tool)
|
||||
assert "ctx" in sig.parameters
|
||||
assert "query" in sig.parameters
|
||||
assert "limit" in sig.parameters
|
||||
assert sig.parameters["limit"].default == 10
|
||||
|
||||
# Check that type hints are preserved
|
||||
type_hints = get_type_hints(example_tool)
|
||||
assert "query" in type_hints
|
||||
assert type_hints["query"] == str
|
||||
assert "limit" in type_hints
|
||||
assert type_hints["limit"] == int
|
||||
assert type_hints["return"] == str
|
||||
|
||||
|
||||
@pytest.mark.asyncio
|
||||
async def test_last_model_retry_soft_fail_normal_execution():
|
||||
"""Test that the decorator doesn't interfere with normal execution."""
|
||||
|
||||
@last_model_retry_soft_fail
|
||||
async def example_tool(_ctx: RunContext, value: str) -> str:
|
||||
"""Example tool."""
|
||||
return f"Result: {value}"
|
||||
|
||||
# Create a mock context
|
||||
class MockContext:
|
||||
"""Fake context for testing."""
|
||||
|
||||
max_retries = 3
|
||||
retries = {}
|
||||
tool_name = "example_tool"
|
||||
|
||||
ctx = MockContext()
|
||||
result = await example_tool(ctx, "test")
|
||||
assert result == "Result: test"
|
||||
|
||||
|
||||
@pytest.mark.asyncio
|
||||
async def test_last_model_retry_soft_fail_handles_retry_exception():
|
||||
"""Test that the decorator handles ModelRetry exceptions correctly."""
|
||||
|
||||
@last_model_retry_soft_fail
|
||||
async def failing_tool(_ctx: RunContext, should_fail: bool) -> str:
|
||||
"""Tool that can raise ModelRetry."""
|
||||
if should_fail:
|
||||
raise ModelRetry("Please retry with different parameters")
|
||||
return "Success"
|
||||
|
||||
# Create a mock context
|
||||
class MockContext:
|
||||
"""Fake context for testing."""
|
||||
|
||||
max_retries = 3
|
||||
retries = {}
|
||||
tool_name = "failing_tool"
|
||||
|
||||
ctx = MockContext()
|
||||
|
||||
# Test when retries haven't been exhausted - should re-raise
|
||||
with pytest.raises(ModelRetry):
|
||||
await failing_tool(ctx, should_fail=True)
|
||||
|
||||
|
||||
@pytest.mark.asyncio
|
||||
async def test_last_model_retry_soft_fail_returns_message_when_max_retries_reached():
|
||||
"""Test that the decorator returns the error message when max retries is reached."""
|
||||
|
||||
@last_model_retry_soft_fail
|
||||
async def failing_tool(_ctx: RunContext, should_fail: bool) -> str:
|
||||
"""Tool that can raise ModelRetry."""
|
||||
if should_fail:
|
||||
raise ModelRetry("Please retry with different parameters.")
|
||||
return "Success"
|
||||
|
||||
# Create a mock context with max retries already reached
|
||||
class MockContext:
|
||||
"""Fake context for testing."""
|
||||
|
||||
max_retries = 3
|
||||
retries = {"failing_tool": 3}
|
||||
tool_name = "failing_tool"
|
||||
|
||||
ctx = MockContext()
|
||||
|
||||
# Test when retries have been exhausted - should return message
|
||||
result = await failing_tool(ctx, should_fail=True)
|
||||
assert result == (
|
||||
"Please retry with different parameters. "
|
||||
"You must explain this to the user and not try to answer based on your knowledge."
|
||||
)
|
||||
|
||||
|
||||
@pytest.mark.asyncio
|
||||
async def test_last_model_retry_soft_fail_returns_message_when_model_cannot_retry():
|
||||
"""Test that the decorator returns the error message when ModelCannotRetry is raised."""
|
||||
|
||||
@last_model_retry_soft_fail
|
||||
async def failing_tool(_ctx: RunContext, should_fail: bool) -> str:
|
||||
"""Tool that can raise ModelRetry."""
|
||||
if should_fail:
|
||||
raise ModelCannotRetry("This is broken duh.")
|
||||
return "Success"
|
||||
|
||||
# Create a mock context with max retries already reached
|
||||
class MockContext:
|
||||
"""Fake context for testing."""
|
||||
|
||||
max_retries = 3
|
||||
retries = {"failing_tool": 3}
|
||||
tool_name = "failing_tool"
|
||||
|
||||
ctx = MockContext()
|
||||
|
||||
# Test when retries have been exhausted - should return message
|
||||
result = await failing_tool(ctx, should_fail=True)
|
||||
assert result == "This is broken duh."
|
||||
File diff suppressed because it is too large
Load Diff
@@ -18,18 +18,28 @@ def get_pydantic_tools_by_name(name: str) -> Tool:
|
||||
tool_dict = {
|
||||
"get_current_weather": Tool(get_current_weather, takes_ctx=False),
|
||||
"web_search_brave": Tool(
|
||||
web_search_brave, takes_ctx=False, prepare=only_if_web_search_enabled
|
||||
web_search_brave,
|
||||
takes_ctx=True,
|
||||
prepare=only_if_web_search_enabled,
|
||||
max_retries=2,
|
||||
),
|
||||
"web_search_brave_with_document_backend": Tool(
|
||||
web_search_brave_with_document_backend,
|
||||
takes_ctx=True,
|
||||
prepare=only_if_web_search_enabled,
|
||||
max_retries=2,
|
||||
),
|
||||
"web_search_tavily": Tool(
|
||||
web_search_tavily, takes_ctx=False, prepare=only_if_web_search_enabled
|
||||
web_search_tavily,
|
||||
takes_ctx=False,
|
||||
prepare=only_if_web_search_enabled,
|
||||
max_retries=2,
|
||||
),
|
||||
"web_search_albert_rag": Tool(
|
||||
web_search_albert_rag, takes_ctx=True, prepare=only_if_web_search_enabled
|
||||
web_search_albert_rag,
|
||||
takes_ctx=True,
|
||||
prepare=only_if_web_search_enabled,
|
||||
max_retries=2,
|
||||
),
|
||||
}
|
||||
|
||||
|
||||
@@ -0,0 +1,23 @@
|
||||
"""Exceptions for tool function retries."""
|
||||
|
||||
from pydantic_ai import ModelRetry
|
||||
|
||||
|
||||
class ModelRetryLast(ModelRetry):
|
||||
"""
|
||||
Same as ModelRetry but also holds the last retry message to return when all attempts failed.
|
||||
"""
|
||||
|
||||
def __init__(self, message: str, last_retry_message: str):
|
||||
"""Initialize ModelRetryLast with message and last retry message."""
|
||||
self.last_retry_message = last_retry_message
|
||||
super().__init__(message)
|
||||
|
||||
|
||||
class ModelCannotRetry(ModelRetry):
|
||||
"""
|
||||
Exception to raise when a tool function cannot be retried.
|
||||
|
||||
We use this exception to signal that the model should not attempt to retry
|
||||
the tool call, typically because the error is not transient or recoverable.
|
||||
"""
|
||||
@@ -0,0 +1,50 @@
|
||||
"""Tool calling utilities for the chat agent."""
|
||||
|
||||
import functools
|
||||
import logging
|
||||
from typing import Any, Callable
|
||||
|
||||
from pydantic_ai import ModelRetry, RunContext
|
||||
|
||||
from chat.tools.exceptions import ModelCannotRetry
|
||||
|
||||
logger = logging.getLogger(__name__)
|
||||
|
||||
|
||||
def last_model_retry_soft_fail(
|
||||
tool_func: Callable[..., Any],
|
||||
) -> Callable[..., Any]:
|
||||
"""
|
||||
Wrap a tool function to handle ModelRetry exceptions.
|
||||
|
||||
If the tool function raises ModelRetry and the maximum number of retries
|
||||
has been reached, a ModelCannotRetry exception is raised instead.
|
||||
|
||||
Args:
|
||||
tool_func: The original tool function to wrap.
|
||||
|
||||
Returns:
|
||||
A wrapped tool function with retry handling.
|
||||
"""
|
||||
|
||||
@functools.wraps(tool_func)
|
||||
async def wrapper(ctx: RunContext, *args, **kwargs) -> Any:
|
||||
try:
|
||||
return await tool_func(ctx, *args, **kwargs)
|
||||
except ModelCannotRetry as exc:
|
||||
return str(exc.message)
|
||||
except ModelRetry as exc:
|
||||
logger.error("Tool '%s' raised ModelRetry: %s", ctx, exc.message)
|
||||
if (ctx.retries.get(ctx.tool_name, 0) + 1) >= ctx.max_retries:
|
||||
logger.error("Max retries reached for tool '%s'.", ctx.tool_name)
|
||||
# A bit of a hack to signal that we cannot retry here, while preventing
|
||||
# the LLM to generate an outdated answer.
|
||||
# We may define a more specific exception later base on ModelRetry which
|
||||
# adds a specific message for this case.
|
||||
return (
|
||||
f"{exc.message} You must explain this to the user and "
|
||||
"not try to answer based on your knowledge."
|
||||
)
|
||||
raise # Re-raise to allow retrying
|
||||
|
||||
return wrapper
|
||||
@@ -1,24 +1,42 @@
|
||||
"""Web search tool using Brave for the chat agent."""
|
||||
|
||||
import asyncio
|
||||
import logging
|
||||
import uuid
|
||||
from concurrent.futures import ThreadPoolExecutor, as_completed
|
||||
from typing import List
|
||||
|
||||
from django.conf import settings
|
||||
from django.core.cache import cache
|
||||
from django.utils.module_loading import import_string
|
||||
from django.utils.text import slugify
|
||||
|
||||
import requests
|
||||
import httpx
|
||||
from asgiref.sync import sync_to_async
|
||||
from pydantic_ai import RunContext, RunUsage
|
||||
from pydantic_ai.exceptions import ModelRetry
|
||||
from pydantic_ai.messages import ToolReturn
|
||||
from trafilatura import extract, fetch_url
|
||||
from trafilatura import extract
|
||||
from trafilatura.meta import reset_caches
|
||||
|
||||
from chat.tools.exceptions import ModelCannotRetry
|
||||
from chat.tools.utils import last_model_retry_soft_fail
|
||||
|
||||
logger = logging.getLogger(__name__)
|
||||
|
||||
|
||||
def llm_summarize(query: str, text: str) -> str:
|
||||
class WebSearchError(Exception):
|
||||
"""Base exception for web search errors."""
|
||||
|
||||
|
||||
class BraveAPIError(WebSearchError):
|
||||
"""Error when calling Brave API."""
|
||||
|
||||
|
||||
class DocumentFetchError(WebSearchError):
|
||||
"""Error when fetching or extracting documents."""
|
||||
|
||||
|
||||
async def llm_summarize_async(query: str, text: str) -> str:
|
||||
"""
|
||||
Summarize the text using the LLM summarization agent.
|
||||
|
||||
@@ -33,7 +51,7 @@ def llm_summarize(query: str, text: str) -> str:
|
||||
prompt = f"""
|
||||
Based on the following request, summarize the following text in a concise manner,
|
||||
focusing on the key points regarding the user request.
|
||||
he result should be up to 30 lines long.
|
||||
The result should be up to 30 lines long.
|
||||
|
||||
<user request>
|
||||
{query}
|
||||
@@ -44,54 +62,87 @@ he result should be up to 30 lines long.
|
||||
</text to summarize>
|
||||
"""
|
||||
|
||||
result = summarization_agent.run_sync(prompt)
|
||||
result = await summarization_agent.run(prompt)
|
||||
return result.output
|
||||
|
||||
|
||||
def _fetch_and_extract(url: str) -> str:
|
||||
"""Fetch and extract text content from the URL."""
|
||||
cache_key = f"web_search_brave:extract:{url}"
|
||||
async def _fetch_url_async(url: str, timeout: int = 30) -> str:
|
||||
"""Fetch URL content asynchronously."""
|
||||
async with httpx.AsyncClient(timeout=timeout, follow_redirects=True) as client:
|
||||
response = await client.get(url)
|
||||
response.raise_for_status()
|
||||
return response.text
|
||||
|
||||
if (document := cache.get(cache_key)) is not None:
|
||||
|
||||
async def _fetch_and_extract_async(url: str) -> str:
|
||||
"""Fetch and extract text content from the URL asynchronously."""
|
||||
cache_key = f"web_search_brave:extract:{slugify(url)}"
|
||||
|
||||
# Check cache first
|
||||
if (document := await cache.aget(cache_key)) is not None:
|
||||
return document
|
||||
|
||||
html = fetch_url(url)
|
||||
document = extract(html, include_comments=False, no_fallback=True) or ""
|
||||
cache.set(cache_key, document, settings.BRAVE_CACHE_TTL)
|
||||
try:
|
||||
# Fetch HTML
|
||||
html = await _fetch_url_async(url, timeout=settings.BRAVE_API_TIMEOUT)
|
||||
|
||||
return document
|
||||
# Extract text in thread pool (trafilatura is CPU-bound)
|
||||
document = await sync_to_async(extract)(html, include_comments=False, no_fallback=True)
|
||||
|
||||
# Cache the result
|
||||
await cache.aset(cache_key, document, settings.BRAVE_CACHE_TTL)
|
||||
return document
|
||||
|
||||
except httpx.HTTPError as e:
|
||||
logger.warning("HTTP error fetching %s: %s", url, e, exc_info=True)
|
||||
raise DocumentFetchError(f"Failed to fetch {url}: {e}") from e
|
||||
except Exception as e:
|
||||
logger.warning("Error extracting content from %s: %s", url, e, exc_info=True)
|
||||
raise DocumentFetchError(f"Failed to extract content from {url}: {e}") from e
|
||||
|
||||
|
||||
def _extract_and_summarize_snippets(query: str, url: str) -> List[str]:
|
||||
async def _extract_and_summarize_snippets_async(query: str, url: str) -> List[str]:
|
||||
"""Fetch, extract and summarize text content from the URL.
|
||||
|
||||
Returns a list of snippets (0 or 1 element, preserving existing behavior).
|
||||
"""
|
||||
# Cache by URL to avoid repeated fetch/extract across calls
|
||||
document = _fetch_and_extract(url)
|
||||
if not document:
|
||||
try:
|
||||
document = await _fetch_and_extract_async(url)
|
||||
if not document:
|
||||
return []
|
||||
|
||||
if not settings.BRAVE_SUMMARIZATION_ENABLED:
|
||||
return [document]
|
||||
|
||||
try:
|
||||
snippet = await llm_summarize_async(query, document)
|
||||
return [snippet] if snippet else []
|
||||
except Exception as e: # pylint: disable=broad-except
|
||||
logger.exception("Summarization failed for %s: %s", url, e)
|
||||
# Fallback to raw document if summarization fails
|
||||
return [document]
|
||||
|
||||
except DocumentFetchError:
|
||||
# Document fetch failed, return empty
|
||||
return []
|
||||
|
||||
if not settings.BRAVE_SUMMARIZATION_ENABLED:
|
||||
return [document]
|
||||
|
||||
async def _fetch_and_store_async(url: str, document_store) -> None:
|
||||
"""Fetch, extract and store text content from the URL in the document store."""
|
||||
|
||||
try:
|
||||
snippet = llm_summarize(query, document)
|
||||
except Exception as e: # pylint: disable=broad-except
|
||||
logger.exception("Summarization failed for %s: %s", url, e)
|
||||
snippet = None
|
||||
document = await _fetch_and_extract_async(url)
|
||||
|
||||
return [snippet] if snippet else []
|
||||
logger.debug("Fetched document: %s", document)
|
||||
|
||||
if document:
|
||||
await document_store.astore_document(url, document)
|
||||
except DocumentFetchError as e:
|
||||
logger.warning("Failed to fetch and store %s: %s", url, e)
|
||||
# Continue with other documents
|
||||
|
||||
|
||||
def _fetch_and_store(url: str, document_store) -> None:
|
||||
"""Fetch, extract and store text content from the URL in the document store."""
|
||||
document = _fetch_and_extract(url)
|
||||
if document:
|
||||
document_store.store_document(url, document)
|
||||
|
||||
|
||||
def _query_brave_api(query: str) -> List[dict]:
|
||||
async def _query_brave_api_async(query: str) -> List[dict]:
|
||||
"""Query the Brave Search API and return the raw results."""
|
||||
url = "https://api.search.brave.com/res/v1/web/search"
|
||||
headers = {
|
||||
@@ -109,14 +160,53 @@ def _query_brave_api(query: str) -> List[dict]:
|
||||
"extra_snippets": settings.BRAVE_SEARCH_EXTRA_SNIPPETS,
|
||||
}
|
||||
params = {k: v for k, v in data.items() if v is not None}
|
||||
response = requests.get(url, headers=headers, params=params, timeout=settings.BRAVE_API_TIMEOUT)
|
||||
response.raise_for_status()
|
||||
|
||||
json_response = response.json()
|
||||
try:
|
||||
async with httpx.AsyncClient(timeout=settings.BRAVE_API_TIMEOUT) as client:
|
||||
response = await client.get(url, headers=headers, params=params)
|
||||
response.raise_for_status()
|
||||
json_response = response.json()
|
||||
|
||||
# See https://api-dashboard.search.brave.com/app/documentation/web-search/responses#Result
|
||||
# & https://api-dashboard.search.brave.com/app/documentation/web-search/responses#SearchResult
|
||||
return json_response.get("web", {}).get("results", [])
|
||||
# https://api-dashboard.search.brave.com/app/documentation/web-search/responses#Result
|
||||
return json_response.get("web", {}).get("results", [])
|
||||
|
||||
except httpx.HTTPStatusError as e:
|
||||
if e.response.status_code == 429:
|
||||
# Rate limit - retryable
|
||||
logger.warning("Brave API rate limited: %s", e)
|
||||
raise ModelRetry(
|
||||
"The search API is rate limited. Please wait a moment and try again."
|
||||
) from e
|
||||
if e.response.status_code >= 500:
|
||||
# Server error - retryable
|
||||
logger.warning("Brave API error: %s", e)
|
||||
raise ModelRetry(
|
||||
"The search service is temporarily unavailable due to a server error. Retrying..."
|
||||
) from e
|
||||
|
||||
# Client error (4xx) - not retryable, stop and inform user
|
||||
logger.error("Brave API client error: %s", e)
|
||||
raise ModelCannotRetry(
|
||||
f"Web search failed with a client error (status {e.response.status_code}). "
|
||||
"You must explain this to the user and not try to answer based on your knowledge."
|
||||
) from e
|
||||
except httpx.TimeoutException as e:
|
||||
# Timeout - retryable
|
||||
logger.warning("Brave API timeout: %s", e)
|
||||
raise ModelRetry("The search request timed out. Retrying with a fresh attempt...") from e
|
||||
except httpx.HTTPError as e:
|
||||
# Other HTTP errors - retryable
|
||||
logger.warning("Brave API connection error: %s", e)
|
||||
raise ModelRetry(
|
||||
f"Connection error while searching the web: {type(e).__name__}. Retrying..."
|
||||
) from e
|
||||
except Exception as e:
|
||||
# Unexpected errors - not retryable, stop completely
|
||||
logger.exception("Unexpected error querying Brave API: %s", e)
|
||||
raise ModelCannotRetry(
|
||||
f"An unexpected error occurred with the search service: {type(e).__name__}. "
|
||||
"You must explain this to the user and not try to answer based on your knowledge."
|
||||
) from e
|
||||
|
||||
|
||||
def format_tool_return(raw_search_results: List[dict]) -> ToolReturn:
|
||||
@@ -140,92 +230,132 @@ def format_tool_return(raw_search_results: List[dict]) -> ToolReturn:
|
||||
)
|
||||
|
||||
|
||||
def web_search_brave(query: str) -> ToolReturn:
|
||||
@last_model_retry_soft_fail
|
||||
async def web_search_brave(_ctx: RunContext, query: str) -> ToolReturn:
|
||||
"""
|
||||
Search the web for up-to-date information
|
||||
|
||||
Args:
|
||||
_ctx (RunContext): The run context, used by the wrapper.
|
||||
query (str): The query to search for.
|
||||
"""
|
||||
raw_search_results = _query_brave_api(query)
|
||||
try:
|
||||
raw_search_results = await _query_brave_api_async(query)
|
||||
|
||||
reset_caches() # Clear trafilatura caches to avoid memory bloat/leaks
|
||||
await sync_to_async(reset_caches)() # Clear trafilatura caches to avoid memory bloat/leaks
|
||||
|
||||
# Parallelize fetch/extract for results that don't include extra_snippets
|
||||
to_process = [
|
||||
(idx, r) for idx, r in enumerate(raw_search_results) if not r.get("extra_snippets")
|
||||
]
|
||||
# Parallelize fetch/extract for results that don't include extra_snippets
|
||||
to_process = [
|
||||
(idx, r) for idx, r in enumerate(raw_search_results) if not r.get("extra_snippets")
|
||||
]
|
||||
|
||||
if to_process:
|
||||
max_workers = min(settings.BRAVE_MAX_WORKERS, len(to_process))
|
||||
if max_workers == 1:
|
||||
# Avoid overhead of ThreadPoolExecutor if only one task
|
||||
for idx, r in to_process:
|
||||
raw_search_results[idx]["extra_snippets"] = _extract_and_summarize_snippets(
|
||||
query, r["url"]
|
||||
)
|
||||
if to_process:
|
||||
# Process all URLs concurrently
|
||||
tasks = [
|
||||
_extract_and_summarize_snippets_async(query, r["url"]) for idx, r in to_process
|
||||
]
|
||||
results = await asyncio.gather(*tasks, return_exceptions=False)
|
||||
|
||||
else:
|
||||
with ThreadPoolExecutor(max_workers=max_workers) as executor:
|
||||
future_map = {
|
||||
executor.submit(_extract_and_summarize_snippets, query, r["url"]): idx
|
||||
for idx, r in to_process
|
||||
}
|
||||
for future in as_completed(future_map):
|
||||
idx = future_map[future]
|
||||
raw_search_results[idx]["extra_snippets"] = future.result()
|
||||
# Update raw_search_results with extracted snippets
|
||||
for (idx, _), snippets in zip(to_process, results, strict=True):
|
||||
raw_search_results[idx]["extra_snippets"] = snippets
|
||||
|
||||
return format_tool_return(raw_search_results)
|
||||
formatted_result = format_tool_return(raw_search_results)
|
||||
|
||||
# Check if we got any valid results
|
||||
if not formatted_result.return_value:
|
||||
raise ModelRetry(
|
||||
"No valid search results were extracted from the web pages. "
|
||||
"Retrying the search to find better sources..."
|
||||
)
|
||||
|
||||
return formatted_result
|
||||
|
||||
except (ModelCannotRetry, ModelRetry):
|
||||
# Re-raise these as-is
|
||||
raise
|
||||
except Exception as exc:
|
||||
# Unexpected error in our code - stop and inform user
|
||||
logger.exception("Unexpected error in web_search_brave: %s", exc)
|
||||
raise ModelCannotRetry(
|
||||
f"An unexpected error occurred during web search: {type(exc).__name__}. "
|
||||
"You must explain this to the user and not try to answer based on your knowledge."
|
||||
) from exc
|
||||
|
||||
|
||||
def web_search_brave_with_document_backend(ctx: RunContext, query: str) -> ToolReturn:
|
||||
@last_model_retry_soft_fail
|
||||
async def web_search_brave_with_document_backend(ctx: RunContext, query: str) -> ToolReturn:
|
||||
"""
|
||||
Search the web for up-to-date information
|
||||
Search the web for up-to-date information using RAG backend
|
||||
|
||||
Args:
|
||||
ctx (RunContext): The run context containing the conversation.
|
||||
query (str): The query to search for.
|
||||
"""
|
||||
raw_search_results = _query_brave_api(query)
|
||||
logger.info("Starting web search with RAG backend for query: %s", query)
|
||||
try:
|
||||
raw_search_results = await _query_brave_api_async(query)
|
||||
|
||||
reset_caches() # Clear trafilatura caches to avoid memory bloat/leaks
|
||||
# Clear trafilatura caches in thread pool to avoid blocking
|
||||
loop = asyncio.get_event_loop()
|
||||
await loop.run_in_executor(None, reset_caches)
|
||||
|
||||
# Store documents in a temporary document store for RAG search
|
||||
document_store_backend = import_string(settings.RAG_DOCUMENT_SEARCH_BACKEND)
|
||||
with document_store_backend.temporary_collection(f"tmp-{uuid.uuid4()}") as document_store:
|
||||
max_workers = min(settings.BRAVE_MAX_WORKERS, len(raw_search_results))
|
||||
if max_workers == 1:
|
||||
for result in raw_search_results:
|
||||
# Fetch and extract document content
|
||||
_fetch_and_store(result["url"], document_store)
|
||||
else:
|
||||
with ThreadPoolExecutor(max_workers=max_workers) as executor:
|
||||
futures = [
|
||||
executor.submit(_fetch_and_store, result["url"], document_store)
|
||||
# Store documents in a temporary document store for RAG search
|
||||
document_store_backend = import_string(settings.RAG_DOCUMENT_SEARCH_BACKEND)
|
||||
|
||||
# Create temporary collection
|
||||
temp_collection_name = f"tmp-{uuid.uuid4()}"
|
||||
try:
|
||||
async with document_store_backend.temporary_collection_async(
|
||||
temp_collection_name
|
||||
) as document_store:
|
||||
# Fetch and store all documents concurrently
|
||||
tasks = [
|
||||
_fetch_and_store_async(result["url"], document_store)
|
||||
for result in raw_search_results
|
||||
]
|
||||
for future in as_completed(futures):
|
||||
try:
|
||||
future.result()
|
||||
except Exception as e: # pylint: disable=broad-except
|
||||
logger.exception("Error fetching/storing document: %s", e)
|
||||
await asyncio.gather(*tasks, return_exceptions=True)
|
||||
|
||||
rag_results = document_store.search(
|
||||
query,
|
||||
results_count=settings.BRAVE_RAG_WEB_SEARCH_CHUNK_NUMBER,
|
||||
)
|
||||
# Perform RAG search
|
||||
rag_results = await document_store.asearch(
|
||||
query,
|
||||
results_count=settings.BRAVE_RAG_WEB_SEARCH_CHUNK_NUMBER,
|
||||
)
|
||||
logger.info("RAG search returned: %s", rag_results)
|
||||
|
||||
ctx.usage += RunUsage(
|
||||
input_tokens=rag_results.usage.prompt_tokens,
|
||||
output_tokens=rag_results.usage.completion_tokens,
|
||||
)
|
||||
ctx.usage += RunUsage(
|
||||
input_tokens=rag_results.usage.prompt_tokens,
|
||||
output_tokens=rag_results.usage.completion_tokens,
|
||||
)
|
||||
|
||||
# Map RAG results back to raw search results to include extra_snippets
|
||||
# Suboptimal O(N^2) but N is small...
|
||||
for rag_result in rag_results.data:
|
||||
for result in raw_search_results:
|
||||
if result["url"] == rag_result.url:
|
||||
result.setdefault("extra_snippets", []).append(rag_result.content)
|
||||
break
|
||||
# Map RAG results back to raw search results to include extra_snippets
|
||||
for rag_result in rag_results.data:
|
||||
for result in raw_search_results:
|
||||
if result["url"] == rag_result.url:
|
||||
result.setdefault("extra_snippets", []).append(rag_result.content)
|
||||
break
|
||||
|
||||
return format_tool_return(raw_search_results)
|
||||
except Exception as exc:
|
||||
logger.exception("Error with document store: %s", exc)
|
||||
raise ModelRetry(
|
||||
f"Document storage temporarily failed: {type(exc).__name__}. "
|
||||
"Retrying the operation..."
|
||||
) from exc
|
||||
|
||||
formatted_result = format_tool_return(raw_search_results)
|
||||
|
||||
# Check if we got any valid results
|
||||
if not formatted_result.return_value:
|
||||
raise ModelRetry("No valid search results were extracted.")
|
||||
|
||||
return formatted_result
|
||||
except (ModelCannotRetry, ModelRetry):
|
||||
# Re-raise these as-is
|
||||
raise
|
||||
except Exception as e:
|
||||
# Unexpected error - stop and inform user
|
||||
logger.exception("Unexpected error in web_search_brave_with_document_backend: %s", e)
|
||||
raise ModelCannotRetry(
|
||||
f"An unexpected error occurred during web search with RAG: {type(e).__name__}. "
|
||||
"You must explain this to the user and not try to answer based on your knowledge."
|
||||
) from e
|
||||
|
||||
@@ -631,9 +631,6 @@ class Base(BraveSettings, Configuration):
|
||||
LLM_DEFAULT_MODEL_HRID = values.Value(
|
||||
"default-model", environ_name="LLM_DEFAULT_MODEL_HRID", environ_prefix=None
|
||||
)
|
||||
LLM_ROUTING_MODEL_HRID = values.Value(
|
||||
"default-routing-model", environ_name="LLM_ROUTING_MODEL_HRID", environ_prefix=None
|
||||
)
|
||||
LLM_SUMMARIZATION_MODEL_HRID = values.Value(
|
||||
"default-summarization-model",
|
||||
environ_name="LLM_SUMMARIZATION_MODEL_HRID",
|
||||
|
||||
@@ -665,6 +665,11 @@ export const Chat = ({
|
||||
{...props}
|
||||
/>
|
||||
),
|
||||
a: ({ children, ...props }) => (
|
||||
<a target="_blank" {...props}>
|
||||
{children}
|
||||
</a>
|
||||
),
|
||||
}}
|
||||
>
|
||||
{message.content}
|
||||
|
||||
@@ -16,7 +16,6 @@ backend:
|
||||
DJANGO_CSRF_TRUSTED_ORIGINS: https://conversations.127.0.0.1.nip.io
|
||||
DJANGO_CONFIGURATION: Feature
|
||||
DJANGO_ALLOWED_HOSTS: conversations.127.0.0.1.nip.io
|
||||
DJANGO_SERVER_TO_SERVER_API_TOKENS: secret-api-key
|
||||
DJANGO_SECRET_KEY: *djangoSecretKey
|
||||
DJANGO_SETTINGS_MODULE: conversations.settings
|
||||
DJANGO_SUPERUSER_PASSWORD: admin
|
||||
|
||||
@@ -19,7 +19,6 @@ backend:
|
||||
DJANGO_CSRF_TRUSTED_ORIGINS: https://conversations.127.0.0.1.nip.io
|
||||
DJANGO_CONFIGURATION: Feature
|
||||
DJANGO_ALLOWED_HOSTS: conversations.127.0.0.1.nip.io
|
||||
DJANGO_SERVER_TO_SERVER_API_TOKENS: secret-api-key
|
||||
DJANGO_SECRET_KEY: *djangoSecretKey
|
||||
DJANGO_SETTINGS_MODULE: conversations.settings
|
||||
DJANGO_SUPERUSER_PASSWORD: admin
|
||||
|
||||
Reference in New Issue
Block a user