Claude 在回答有關文件的問題時可以提供詳細的引用,協助您追蹤和驗證每個回應背後的來源。
所有現行模型都支援引用功能。
以下範例展示如何透過 Messages API 在純文字文件上啟用引用:
client = anthropic.Anthropic()
response = client.messages.create(
model="claude-opus-5",
max_tokens=1024,
messages=[
{
"role": "user",
"content": [
{
"type": "document",
"source": {
"type": "text",
"media_type": "text/plain",
"data": "The grass is green. The sky is blue.",
},
"title": "My Document",
"context": "This is a trustworthy document.",
"citations": {"enabled": True},
},
{"type": "text", "text": "What color is the grass and sky?"},
],
}
],
)
print(response)透過以下步驟將引用功能與 Claude 整合:
提供文件並啟用引用
文件處理
Claude 提供帶引用的回應
source 內容中找到的文字可以被引用。title 和 context 是可選欄位,會傳遞給模型,但不會用於被引用的內容。title 的長度有限制,因此 context 欄位適合用於以文字或字串化 JSON 的形式儲存文件中繼資料。content 清單從 0 開始編號,結束索引為不包含。cited_text 欄位是為了方便而提供,不會計入輸出 token。cited_text 也不會計入輸入 token。引用功能可與其他 API 功能搭配使用,包括提示快取、token 計數和批次處理。
引用和提示快取可以有效地一起使用。
回應中產生的引用區塊無法直接快取,但它們所參照的來源文件可以快取。為了最佳化效能,請將 cache_control 套用至您的頂層文件內容區塊。
client = anthropic.Anthropic()
# 長篇文件內容(例如技術文件)
long_document = (
"This is a very long document with thousands of words..." + " ... " * 1000
) # Minimum cacheable length
response = client.messages.create(
model="claude-opus-5",
max_tokens=1024,
messages=[
{
"role": "user",
"content": [
{
"type": "document",
"source": {
"type": "text",
"media_type": "text/plain",
"data": long_document,
},
"citations": {"enabled": True},
"cache_control": {
"type": "ephemeral"
}, # Cache the document content
},
{
"type": "text",
"text": "What does this document say about API features?",
},
],
}
],
)
print(response)在此範例中:
cache_control 進行快取。引用支援三種文件類型。文件可以直接在訊息中提供(base64、文字或 URL),或透過 Files API 上傳並透過 file_id 參照:
| 類型 | 最適合 | 分塊方式 | 引用格式 |
|---|---|---|---|
| 純文字 | 簡單的文字文件、散文 | 句子 | 字元索引(從 0 開始編號) |
| 含文字內容的 PDF 檔案 | 句子 | 頁碼(從 1 開始編號) | |
| 自訂內容 | 清單、逐字稿、特殊格式、更細粒度的引用 | 無額外分塊 | 區塊索引(從 0 開始編號) |
純文字文件會自動分塊為句子。您可以內嵌提供它們,或透過其 file_id 參照:
本頁頂部的入門範例展示了每個 SDK 中完整的純文字請求。文件區塊使用 text 來源:
{
"type": "document",
"source": {
"type": "text",
"media_type": "text/plain",
"data": "Plain text content..."
},
"title": "Document Title",
"context": "Context about the document that will not be cited from",
"citations": { "enabled": true }
}PDF 文件可以以 base64 編碼資料、URL 或 file_id 的形式提供。PDF 文字會被擷取並分塊為句子。由於尚不支援圖片引用,因此屬於文件掃描檔且不包含可擷取文字的 PDF 無法被引用。
client = anthropic.Anthropic()
pdf_base64 = base64.standard_b64encode(
pathlib.Path("/path/to/document.pdf").read_bytes()
).decode()
response = client.messages.create(
model="claude-opus-5",
max_tokens=1024,
messages=[
{
"role": "user",
"content": [
{
"type": "document",
"source": {
"type": "base64",
"media_type": "application/pdf",
"data": pdf_base64,
},
"title": "Document Title",
"context": "Context about the document that will not be cited from",
"citations": {"enabled": True},
},
{"type": "text", "text": "Summarize this document."},
],
}
],
)
print(response)自訂內容文件讓您可以控制引用粒度。不會進行額外的分塊,區塊會根據所提供的內容區塊提供給模型。
client = anthropic.Anthropic()
response = client.messages.create(
model="claude-opus-5",
max_tokens=1024,
messages=[
{
"role": "user",
"content": [
{
"type": "document",
"source": {
"type": "content",
"content": [
{"type": "text", "text": "First chunk"},
{"type": "text", "text": "Second chunk"},
],
},
"title": "Document Title",
"context": "Context about the document that will not be cited from",
"citations": {"enabled": True},
},
{"type": "text", "text": "Summarize this document."},
],
}
],
)
print(response)啟用引用時,回應會包含多個帶有引用的文字區塊:
{
"content": [
{"type": "text", "text": "According to the document, "},
{
"type": "text",
"text": "the grass is green",
"citations": [
{
"type": "char_location",
"cited_text": "The grass is green.",
"document_index": 0,
"document_title": "Example Document",
"start_char_index": 0,
"end_char_index": 20,
}
],
},
{"type": "text", "text": " and "},
{
"type": "text",
"text": "the sky is blue",
"citations": [
{
"type": "char_location",
"cited_text": "The sky is blue.",
"document_index": 0,
"document_title": "Example Document",
"start_char_index": 20,
"end_char_index": 36,
}
],
},
{
"type": "text",
"text": ". Information from page 5 states that ",
},
{
"type": "text",
"text": "water is essential",
"citations": [
{
"type": "page_location",
"cited_text": "Water is essential for life.",
"document_index": 1,
"document_title": "PDF Document",
"start_page_number": 5,
"end_page_number": 6,
}
],
},
{
"type": "text",
"text": ". The custom document mentions ",
},
{
"type": "text",
"text": "important findings",
"citations": [
{
"type": "content_block_location",
"cited_text": "These are important findings.",
"document_index": 2,
"document_title": "Custom Content Document",
"start_block_index": 0,
"end_block_index": 1,
}
],
},
]
}對於串流回應,引用會以 citations_delta 差異類型的形式出現在 content_block_delta 事件中。每個差異包含一個要新增至目前 text 內容區塊上 citations 清單的單一引用。
在處理文字差異的同時處理 citations_delta 差異類型,以便在串流時呈現帶引用的回應。
將來自您 RAG 管線的搜尋結果作為具有內建引用支援的一級內容區塊傳遞。
了解 Claude 如何從 PDF 擷取文字,以及基於頁面的引用如何對應回您的來源檔案。
上傳文件一次,並在多個引用請求中透過 file_id 參照它們。
| Supported platforms |
|
|---|
Was this page helpful?