mirror of
https://github.com/MindWorkAI/AI-Studio.git
synced 2026-10-11 22:53:48 +00:00
3.3 KiB
3.3 KiB
AGENTS.md
These instructions cover the indexing pipeline in app/MindWork AI Studio/Tools/Services/Indexing/ and
DataSourceEmbeddingService. They add to the AGENTS.md in the repository root, which applies here as
well.
Indexed data sources
The parts:
IIndexedDataSource(Settings/) - what indexing needs to know about a data source: confidence level, embedding provider and chunk settings.IDataSourceBaseis what every data source has. ImplementIDataSourceon top only when classic RAG, Semantic Search and the agents should see the data source. A data source kept in a list of its own implementsIIndexedDataSourcealone, and the compiler keeps it out ofDataSources.IIndexedSourceIndexer- one per kind of data source.Supportsclaims the data sources,ProcessAsyncfinds and reads the documents of one run, andTrackChanges/StopTrackingnotice changes on their own.FileSourceIndexeris the reference: a file system watcher per data source, fingerprints over path, size and write time.IndexedRunContext- one prepared run: both stores, the embedding provider, the manifest and the collection.IndexDocumentAsyncembeds and stores one document; the cleanup methods remove what a failed attempt left behind.EmbeddingDocument- one document: its key, its index row, its display name, and how to read its chunks.DocumentRunProgress- counts the documents, records indexed and failed ones in the stores, publishes the status and completes the run.TextChunker- cuts text into chunks the embedding provider accepts. Pick one of its strategies; do not write a chunker of your own.
To add a kind of data source:
- Write its indexer in
Tools/Services/Indexing/and create it inDataSourceEmbeddingService.CreateIndexers, which hands every indexer the sameTextChunker. - Gate it in
IsSupportedIndexedSource, behind a preview feature of its own while it is new. - When the data source is not kept in
DataSources, add its list toGetConfiguredIndexedSources. Every lookup by id and every pass over all data sources goes through it: the startup hash check,QueueAllInternalDataSourcesAsyncandRefreshWatchers. - Keep whatever the kind has to remember beyond its documents in tables of its own in the index store, added
by an EF Core migration (see
app/MindWork AI Studio/Tools/Databases/AGENTS.md). - Report every status through
DocumentRunProgress, so all rows of the embedding page behave alike.
Rules which are easy to break:
- A document key is not a path. Only files use their full path as the key. Never pass a key through the
PathAPIs: on Windows,Path.GetFullPathreads a key likemail:…as a file with an alternate data stream. - Ids and signature are pinned. The formats in
IndexedDocumentIdsand the embedding signature (DataSourceEmbeddingService.BuildEmbeddingSignature) are fixed by tests, because every stored chunk and every index depends on them. When the metadata stored next to a chunk changes, raiseCHUNK_METADATA_VERSIONdeliberately: that rebuilds every index. - The service decides when, the indexer decides how. Whether changes are tracked at all depends on the automatic refresh setting and the startup hash check, and only the embedding service decides that.