Skip to main content

Data Sources and Datasets

The Console Data module manages three different concepts: relational data access, GeniSpace-managed datasets, and platform-provided data. Choose the data type before configuring an agent, workflow, or Workbench binding.

Open the Data module​

Open Console → Data and choose a tab:

TabPurpose
Data SourcesConnect databases and define controlled SQL-based operations
DatasetsCreate managed collections for structured, full-text, and vector search
Platform DataView data resources exposed by the platform or installed applications

Datasources​

A Datasource connects to a relational database and exposes a defined operation rather than unrestricted database access.

Typical workflow:

  1. Configure a database connection with secure credentials.
  2. Create a Datasource operation against that connection.
  3. Define SQL, parameters, and the allowed read or write behavior.
  4. Test with non-production values.
  5. Authorize the Datasource for a workflow, agent, or Workbench component.

Use separate operations for read, create, update, and delete. Do not grant a general write operation when the user needs only a lookup.

Datasets​

A Dataset stores schema-defined records managed by GeniSpace. It can support:

  • Exact primary-key and filter queries
  • Record count
  • Insert, update, and delete
  • Full-text search on configured text fields
  • Vector search over configured vectorized source fields

Datasets are separate from agent knowledge-base storage. Binding a Dataset to an agent does not make it part of the knowledge base.

Create a Dataset​

  1. Open Data → Datasets.
  2. Select Create Dataset.
  3. Enter the name and description.
  4. Define business fields and a primary key.
  5. Enable vector search or full-text search only when needed.
  6. For vector search, mark the text fields whose content should contribute to retrieval.
  7. For full-text search, enable eligible text fields and select the analyzer when offered.
  8. Review the generated index fields and create the Dataset.

Field names become an API and agent-tool contract. Use stable, descriptive names and do not rename fields casually after applications depend on them.

Vectorization and embedding model​

When vector search is enabled, GeniSpace sends the selected source-field content through the platform model gateway and stores the generated vector. Query text is embedded with the same model bound to that Dataset.

  • Ordinary users provide text, not precomputed vectors.
  • The Dataset records its actual embedding configuration.
  • The embedding model and vector dimension must remain consistent for writes and searches.
  • Text Datasets use the edition's configured Dataset embedding model.
  • A Dataset that requires image or video semantics must be explicitly configured for a supported multimodal model.

Changing a model or source-field projection can require re-vectorizing existing records. Coordinate such changes with an administrator.

NeedOperation
Known IDs or exact field conditionsQuery
Total number of recordsCount
Literal words, names, codes, or phrasesFull-text search
Similar meaning, experience, capability, or cross-language conceptVector search

Vector Top-K returns the nearest records. It does not prove that a record satisfies a mandatory business condition. For example, retrieve candidates semantically, then verify required certification, location, or education fields with structured data.

Dataset details and preview​

The Dataset page shows schema, index/search settings, statistics, and a record preview. Display fields can be adjusted without changing the underlying schema. Vector values may be abbreviated in the preview.

Statistics may be temporarily unavailable while the collection is loading or the vector database is unavailable. A statistics error does not by itself mean the API request was unauthorized.

API Playground​

Open a Dataset and select API Playground. It generates schema-aware request bodies and code examples for the selected Dataset.

OperationEndpoint suffixMain request fields
Insert/insertdata array of records
Update/updatefilter, update_data
Query/queryids or filter, limit, offset, outputFields
Count/countoptional filter depending on the contract
Delete/deletefilter; requires confirmation in Playground
Vector search/searchtext, limit, optional output fields/filter
Full-text search/full-text-searchquery data, sparse index field, limit, output fields

All operations are sent as POST requests under:

https://api.genispace.ai/api/datasets/{datasetId}/data/{operation}

The Playground provides cURL, Python, and JavaScript examples. Replace YOUR_TOKEN with a valid API key or access token and never publish that credential.

Insert records​

{
"data": [
{
"name": "Example record",
"description": "Content used for retrieval"
}
]
}

Vectorized fields are embedded per record. A multi-record insert may use a batch database operation, but each record's semantic content remains an independent embedding input.

Type-safe ID query​

Use ids when primary keys are known:

{
"ids": ["record-001", "record-002"],
"limit": 10,
"offset": 0,
"outputFields": ["name", "description"]
}

Use strings for string primary keys and numbers for numeric primary keys. The API constructs the database expression safely instead of requiring the caller to concatenate an IN (...) expression.

Update records​

{
"filter": "id == \"record-001\"",
"update_data": {
"description": "Updated description"
}
}

Updating a vectorized source field regenerates its vector.

Delete records​

Delete permanently removes every record matching the filter. Test the same filter with Query first and use the Playground confirmation.

Filter and field rules​

  • Use only fields declared by the Dataset schema.
  • Request only existing outputFields; an unknown field causes the vector database to reject the query.
  • Use the expression syntax shown by the Playground. Do not invent SQL functions for Dataset filters.
  • Prefer ids for primary-key batches instead of building an IN expression yourself.
  • Keep mandatory structured conditions separate from semantic query text.

Use a Dataset with an agent​

  1. Edit the agent in Console.
  2. Authorize the required Dataset.
  3. Enable the appropriate dataset tools.
  4. Give the agent an accurate description of the schema and business meaning.
  5. Test exact, full-text, and semantic cases separately.
  6. Verify empty results and tool errors are reported honestly.

Troubleshooting​

ProblemCheck
Vector search returns no recordsCollection data, vectorized source fields, embedding status, Dataset-bound model, and query filter
field ... not existRemove unknown output/filter fields and use the Dataset schema
Filter parse errorUse supported comparison syntax or the ids field for primary keys
Full-text results are too broadVerify the configured source field, analyzer, and query terms
Insert succeeds but semantic search does notCheck vector generation status and whether the intended source fields are marked vectorized
Statistics failCheck collection readiness and vector-database health; retry after the service recovers