Pushing Files to a Custom Source
Upload PDFs, Word documents, and other files to a custom source, and let Guru extract, tag, and sync their text.
If the data you want in Guru lives in files, such as PDFs, Word documents, or other office files, you don't have to extract the text yourself. Upload each file to a custom source and Guru extracts its text, stores it as a searchable record, and keeps it in the same open/push/close sync as any other record.
This guide extends Creating a Custom Source. It assumes you know how to create a source, open a sync, and close it. The examples build a source of company policy documents: one RECORD object type, filterable by file extension, modified date, and department.
Endpoints
All endpoints below are relative to the API root https://api.getguru.com/api/v1/ and use basic auth with your Guru user and API token.
| Endpoint | Purpose |
|---|---|
PUT /sources/{sourceId}/objecttypes/{objectTypeId}/objects/{externalId}/file | Upload a file. Guru extracts its text and creates or updates the record. |
GET /sources/{sourceId}/objecttypes/{objectTypeId}/objects/{externalId}/metadata | Read a record's stored metadata, including the checksum you sent with it. |
PUT /sources/{sourceId}/objecttypes/{objectTypeId}/objects/{externalId} | Mark an unchanged record as part of the current sync without re-uploading it. |
Folder trees use a separate set of endpoints, covered in Organizing a Custom Source into Folders.
Step 1: Create the source
Files belong in a RECORD object type. Records of a RECORD type have a fixed shape (title, url, and content), so the object type needs no templates: Guru fills in all three from the upload. See Structured and record object types for how RECORD differs from STRUCTURED.
Files usually come with metadata worth filtering on, such as their type or when they last changed. Tags carry that metadata. You declare each tag on the object type with a tagConfigs entry, expose it as a filter with a TAG facet, and set its value on every upload with a header (Step 3).
Tags take the place fields play for JSON records. A JSON push puts filter values in the body, where a FIELD facet reads them through a contentSelector. A file upload has no body of yours to read: Guru builds the record's title, url, and content from the file itself. So for files, use tags for anything people should filter on. See Fields or tags for the comparison.
Request
curl -X POST "https://api.getguru.com/api/v1/sources" \
-u $GURU_USER:$GURU_TOKEN \
-H "Content-Type: application/json" \
-d '{
"definition": {
"type": "CUSTOM"
},
"externalId": "acme-policies",
"config": {
"type": "CUSTOM",
"name": "Acme Policies",
"trackStatus": true,
"specification": {
"application": {
"objectTypes": [
{
"name": "Document",
"externalId": "document",
"sourceDataType": "RECORD",
"fields": [
{"name": "title", "contentSelector": "/title", "dataType": "TEXT"},
{"name": "url", "contentSelector": "/url", "dataType": "TEXT"},
{"name": "content", "contentSelector": "/content", "dataType": "TEXT"}
],
"tagConfigs": [
{"type": "SIMPLE", "name": "Extension", "externalId": "extension", "dataType": "TEXT", "allowMultipleValues": false},
{"type": "SIMPLE", "name": "ModifiedDate", "externalId": "modifiedDate", "dataType": "DATE_TIME", "allowMultipleValues": false},
{"type": "SIMPLE", "name": "Department", "externalId": "department", "dataType": "TEXT", "allowMultipleValues": true}
],
"facets": [
{"type": "TAG", "name": "Extension", "tagConfig": {"externalId": "extension"}, "allowMultipleValues": false},
{"type": "TAG", "name": "Modified Date", "tagConfig": {"externalId": "modifiedDate"}, "targetType": "MODIFIED_DATE", "allowMultipleValues": false},
{"type": "TAG", "name": "Department", "tagConfig": {"externalId": "department"}, "allowMultipleValues": true}
]
}
]
}
}
}
}'As with any custom source, store the returned source id and the object type id from sourceObjectTypes. Every call below uses both. Keep trackStatus: true in the request: the upload endpoint returns 404 on a source created without it.
Tag configs
Each tagConfigs entry declares one tag that records of this type can carry.
| Field | Description |
|---|---|
type | SIMPLE for a tag that takes plain values, or HIERARCHICAL for a folder tree (see Organizing a Custom Source into Folders). |
name | The tag's display name. |
externalId | Your identifier for the tag. Uploads and facets refer to the tag by this value. |
dataType | The type of the tag's values, such as TEXT, DATE_TIME, or LONG. |
allowMultipleValues | Whether one record can carry more than one value for this tag. |
Tag facets
A facet with type TAG exposes a tag as a filter in Guru's search and in Knowledge Agent filters. It is the tag equivalent of the FIELD facets in Creating a Custom Source.
| Field | Description |
|---|---|
type | TAG to facet on a tag's values. |
name | The facet's display name in Guru's filters. |
tagConfig.externalId | The externalId of the tag to facet on. It must exist in tagConfigs. |
targetType | Optional. CREATED_DATE or MODIFIED_DATE maps a DATE_TIME tag onto Guru's built-in date filters. Leave it out for any other tag. |
allowMultipleValues | Whether a single record can have more than one value for this facet. |
Step 2: Open the sync
Open a sync on the object type exactly as in Creating a Custom Source: START_INITIAL for the first sync, START after that. Read the sync number from objectType.currentSyncNumber in the response and pass it on every upload and touch in the steps below.
Step 3: Upload files
Send each file as multipart/form-data, with the file in a part named file. $RECORD_EXTERNAL_ID in the path is your system's unique id for the file, such as doc-1001, and the query parameters tie the upload to the open sync and give Guru what the file itself can't: its link back to the source system and a checksum for skipping it next time.
Request
curl -X PUT "https://api.getguru.com/api/v1/sources/$SOURCE_ID/objecttypes/$OBJECT_TYPE_ID/objects/$RECORD_EXTERNAL_ID/file?syncNumber=$SYNC_NUMBER&externalChecksum=d4e5f6&externalUrl=https%3A%2F%2Fintranet.example.com%2Fpolicies%2Freturns" \
-u $GURU_USER:$GURU_TOKEN \
-H "X-Guru-Object-Tag: extension:pdf" \
-H "X-Guru-Object-Tag: modifiedDate:2026-09-01T12:00:00Z" \
-H "X-Guru-Object-Tag: department:Support" \
-H "X-Guru-Object-Tag: department:Finance" \
-F "[email protected];type=application/pdf"-F sets the Content-Type header to multipart/form-data for you.
Query parameters
| Parameter | Description |
|---|---|
syncNumber | The sync number from Step 2. Ties the upload to the open sync. Leave it out to upload outside a sync (see Push outside a sync). |
externalUrl | The click-through link shown with search results and citations. Point it at the file in its source system. URL-encode it. |
externalChecksum | Optional. Any string that changes when the file changes, such as the hash or version your source system reports. Guru stores it and uses it to skip unchanged files (Step 4). |
Headers
| Header | Description |
|---|---|
X-Guru-Object-Tag | Optional. One tag value in the form <tagExternalId>:<value>. Repeat the header for each value, including several values of the same tag. The tag must exist in the object type's tagConfigs, or the request fails with a 404. |
Response
A new record returns 201 Created. Uploading to an externalId that already exists, including one removed by an earlier sync, returns 200 OK and increments version.
{
"version": 1,
"id": "91f30634-e8ea-402c-bc13-3e6322f7b9de",
"objectType": {
"name": "Document",
"id": "687b0694-2f0a-442b-a651-54b3562646b2",
"currentSyncNumber": 1
},
"externalId": "doc-1001",
"createdDate": "2026-09-22T23:56:32.934+0000",
"modifiedDate": "2026-09-22T23:56:32.974+0000",
"externalChecksum": "d4e5f6",
"modifiedBy": {
"email": "[email protected]",
...
}
}What Guru stores
For each upload, Guru builds the record from the file and your parameters:
| Record field | Where it comes from |
|---|---|
title | The file's name, including its extension. Guru drops accents and other non-ASCII characters, then replaces each run of characters other than letters, digits, -, _, ., and ' with a single underscore, so Q3 Report (Final).pdf becomes Q3_Report_Final_.pdf. |
url | The externalUrl parameter. |
content | The text Guru extracts from the file. This is what search and answers draw on. |
Guru extracts text from these formats:
| Kind | Formats |
|---|---|
| Documents | PDF (with a text layer), Word (.docx, .doc), OpenDocument text (.odt), RTF, EPUB |
| Presentations | PowerPoint (.pptx) |
| Spreadsheets | Excel (.xlsx), CSV |
| Web and text | HTML, Markdown, plain text, JSON, XML |
Guru detects the format from the file itself, so you can send any file with curl's default content type.
Guru doesn't run OCR on uploaded files. Images, such as PNG or JPEG screenshots and photos, and scanned PDFs that contain only images of pages produce no text. Guru keeps roughly the first million characters of extracted text and drops the rest.
A file Guru can't read is still stored and returns a success status, but with empty content, so it never matches a search. Check the formats above before you sync a library, so you don't upload files Guru can't use.
The file name becomes the titleThe file's name becomes the record's title exactly as sent, so
Returns and Refunds Policy.pdfis displayed asReturns_and_Refunds_Policy.pdf. Name the file part with the title you want people to see. If you need spaces or other characters in the title, extract the text yourself and push the record as JSON with the/contentendpoint instead (see Structured and record object types).
Tags are replaced on every upload, not merged. An upload that sends only extension:pdf leaves the record with that one tag, even if an earlier upload set others. Send the full set of tags every time.
The X-Guru-Object-Tag header works the same way on JSON pushes to the /content endpoint, so a RECORD object type can use tags whichever way you send its records.
Step 4: Skip unchanged files
Re-uploading every file on every full sync is slow for large libraries. If you send an externalChecksum with each upload, later syncs can check whether a file changed and skip the upload when it didn't. The record still has to be marked as part of the sync, or closing with COMPLETE removes it.
For each file on a later sync:
- Fetch the record's metadata and compare its
externalChecksumwith the file's current checksum. - If they match, touch the record to mark it as part of this sync.
- If they don't match, or the record doesn't exist, upload the file as in Step 3.
Read the stored checksum
curl "https://api.getguru.com/api/v1/sources/$SOURCE_ID/objecttypes/$OBJECT_TYPE_ID/objects/$RECORD_EXTERNAL_ID/metadata" \
-u $GURU_USER:$GURU_TOKENThe response has the same shape as the upload response, including externalChecksum. An externalId that was never uploaded returns a 404.
A record removed by an earlier sync still returns metadata, but without an externalChecksum. Comparing checksums therefore treats it as changed, and the upload in step 3 restores it.
Touch an unchanged record
curl -X PUT "https://api.getguru.com/api/v1/sources/$SOURCE_ID/objecttypes/$OBJECT_TYPE_ID/objects/$RECORD_EXTERNAL_ID?syncNumber=$SYNC_NUMBER" \
-u $GURU_USER:$GURU_TOKEN \
-H 'If-Match: "d4e5f6"'The touch sends no body. It adds the record to the current sync and changes nothing else.
| Status | Meaning |
|---|---|
204 No Content | The record is part of this sync. |
412 Precondition Failed | The If-Match checksum doesn't match the stored one. Upload the file instead. |
404 Not Found | No current record has this externalId, including one removed by an earlier sync. Upload the file instead. |
The If-Match header is optional. Without it, the touch always succeeds for an existing record, so send it whenever you want the server to double-check the comparison you made.
You can skip the metadata callYou can also skip the metadata call and just upload every file with its checksum. When the checksum matches the stored one, Guru doesn't re-extract the file: it touches the record and returns
200 OKwith the current metadata. This saves processing but not bandwidth, since the file is still sent. It also ignores any changedX-Guru-Object-Tagheaders orexternalUrlon that upload, so change the checksum whenever you need those updated.
Step 5: Close the sync
Close the sync with COMPLETE or COMPLETE_INCREMENTAL, exactly as in Creating a Custom Source. On a full sync, every record you uploaded or touched with the current sync number stays, and every other record is removed. If the run can't finish, close it with FAIL instead, which removes nothing. There is a short delay before new files are searchable.
Organize files into folders
If your files live in a folder tree, tag each one with its folder so people can filter by folder and scope Knowledge Agents to part of the tree. Folders use a HIERARCHICAL tag with its own sync; see Organizing a Custom Source into Folders.
Attach a file to a structured record
The same endpoint can fill one field of a STRUCTURED record with a file's text, for example a support ticket whose searchable body includes an attached PDF. This takes two things the RECORD flow doesn't.
First, the object type must declare "dataType": "JSON" at the object type level, next to sourceDataType:
{
"name": "Ticket",
"sourceDataType": "STRUCTURED",
"dataType": "JSON",
"fields": [
{"name": "title", "contentSelector": "/title", "dataType": "TEXT"},
{"name": "url", "contentSelector": "/url", "dataType": "TEXT"},
{"name": "attachment", "contentSelector": "/attachment", "dataType": "TEXT"}
],
"templates": {
"titleTemplate": "${f_title}",
"externalUrlTemplate": "${f_url}",
"searchTemplate": "${f_title}\n${f_attachment!}"
}
}Second, the upload has exactly two parts. The file part holds the record's JSON, the same body you would push to /content. The other part holds the file, and its name is the id of the field the extracted text goes into, taken from sourceObjectTypes[].fields[] in the create response:
curl -X PUT "https://api.getguru.com/api/v1/sources/$SOURCE_ID/objecttypes/$OBJECT_TYPE_ID/objects/$RECORD_EXTERNAL_ID/file?syncNumber=$SYNC_NUMBER" \
-u $GURU_USER:$GURU_TOKEN \
-F "[email protected];type=application/json" \
-F "[email protected];type=application/pdf"Guru parses the JSON, extracts the PDF's text, and stores it at the field's contentSelector (here, /attachment) before building the record from your templates.
| Error | Cause |
|---|---|
400 Object DataType must be of type JSON | The object type has no "dataType": "JSON". This is set at creation. |
400 Expected exactly 2 parts in the multipart request | The request has only the file part, or more than two parts. |
400 ID is invalid | The second part is named with the field's name instead of its id. |
Updated 1 day ago

