Split AGENTS.md into area guides so Codex reads it in full again (#1046)

This commit is contained in:
Thorsten Sommer authored and GitHub committed 2026-10-11 11:22:40 +02:00
1 parent 8badadb35f
commit 3a1f586f76
17 files changed
+435 -357

No files matched your search

+29 -357
View File
@@ -1,6 +1,7 @@
# AGENTS.md
This file provides guidance to Claude Code (claude.ai/code) when working with code in this repository.
This file provides guidance to coding agents, such as Claude Code and Codex, when working with code in this
repository.
## Incremental implementation workflow
@@ -13,6 +14,29 @@ a time. After each item:
4. Stop and wait until the developer has reviewed and committed the changes before continuing.
5. Never push the changes; the developer performs all pushes.
## Area guides
Some areas of the code have instructions of their own, in an `AGENTS.md` next to the code. Claude Code
loads such a guide by itself once it works in that directory; other agents, Codex among them, do not.
Therefore, before you plan, read, or change anything in one of these areas, read its guide first:
| Before you work on … | read first |
|---|---|
| any C#, Razor, or Lua code, or anything else under `app/` | `app/AGENTS.md` |
| anything under `runtime/` | `runtime/AGENTS.md` |
| a release, or crediting a contributor in a changelog entry or after a merge | `app/Build/AGENTS.md` |
| anything under `app/MindWork AI Studio/Tools/PluginSystem/`, or a new capability of configuration plugins | `app/MindWork AI Studio/Tools/PluginSystem/AGENTS.md` |
| anything under `app/MindWork AI Studio/Tools/ToolCallingSystem/` | `app/MindWork AI Studio/Tools/ToolCallingSystem/AGENTS.md` |
| anything under `app/MindWork AI Studio/Models/` or `app/Tests/Models/` | `app/MindWork AI Studio/Models/AGENTS.md` |
| anything under `app/MindWork AI Studio/Tools/Services/Indexing/`, or `DataSourceEmbeddingService` | `app/MindWork AI Studio/Tools/Services/Indexing/AGENTS.md` |
| anything under `app/MindWork AI Studio/Tools/Databases/` | `app/MindWork AI Studio/Tools/Databases/AGENTS.md` |
This file keeps only what applies to the whole repository. `app/AGENTS.md` keeps what applies to all of the
.NET app, such as the rules for code which calls into one of the areas below it. Add instructions about a
single area to its guide instead. A new guide is an `AGENTS.md` plus a `CLAUDE.md` holding only
`@AGENTS.md`, listed in the table above. Never place one below `app/MindWork AI Studio/wwwroot/` or
`app/MindWork AI Studio/Plugins/`: both are embedded into the app and shipped to every user.
## Working for an external contributor
When you work for someone outside the core team, `CONTRIBUTING.md` applies in addition to this file. In
@@ -105,344 +129,20 @@ through the IDE for the same reason they build there:
mcp__rider__execute_terminal_command command: "cd app/Tests && dotnet test"
```
An assembly-wide `[SetUpFixture]` in `app/Tests/TestHost.cs` fills the static application state that
the app itself only fills while starting up, `Program.LOGGER_FACTORY` above all. Types that
initialize a static logger from it — `Settings.Provider` among them — otherwise die in their type
initializer before the first assertion. Prefer writing new code so that it does not reach for such
statics at all.
The Rust tests run with `cargo test` in `runtime/`, through the `rustrover` MCP server.
## Architecture Details
### Rust Runtime (`runtime/`)
**Entry point:** `runtime/src/main.rs`
Key modules:
- `app_window.rs` - Tauri window management, updater integration
- `dotnet.rs` - Launches and manages the .NET sidecar process
- `runtime_api.rs` - Axum-based HTTPS API for .NET ↔ Rust communication
- `certificate.rs` - Generates self-signed TLS certificates for secure IPC
- `secret.rs` - Secure secret storage using OS keyring (Keychain/Credential Manager)
- `clipboard.rs` - Cross-platform clipboard operations
- `file_data.rs` - File processing for RAG (extracts text from PDF, DOCX, XLSX, PPTX, etc.)
- `encryption.rs` - AES-256-CBC encryption for sensitive data
- `pandoc.rs` - Integration with Pandoc for document conversion
- `log.rs` - Logging infrastructure using `flexi_logger`
**Every runtime API route requires the API token.** `require_api_token` in `runtime_api.rs` checks it
for all routes at once, so handlers take no `APIToken` argument. Register a new route in
`create_router`, before that call: `route_layer` protects only the routes registered before it, and
a route added afterward would be open to every process on the machine. The test
`every_route_requires_the_api_token` reads the routes from `create_router` and fails for any route
which answers without the token.
**Runtime API handlers never block.** All calls of the .NET app share one HTTP/2 connection, and the
task driving it may wait on exactly the Tokio worker which a blocking handler occupies, so a single
blocking handler can hold up the whole app. Therefore:
- File and disk access, the OS keyring, waiting for a `std::sync::Mutex`, and CPU-bound work such as
loading a tokenizer or scanning text run in `tokio::task::spawn_blocking`. Follow
`run_qdrant_edge_request` in `qdrant_edge_database.rs`, `prepare_image` and `prepare_image_sync` in
`image.rs`, or `sanitize_batch` in `prompt_injection/api.rs`.
- When an API offers a callback, as the file dialogs do, await the callback instead of calling the
blocking variant; see `await_dialog` in `file_actions.rs`.
- Short, bounded work, such as a single metadata lookup or writing a log line, may stay on the worker.
- When the blocking task fails, answer with an error, never with an empty value the app could take
for a valid answer.
- Never keep a `std::sync::MutexGuard` alive across an `.await`, not even with an explicit `drop`
before it: the future of the handler is then no longer `Send`, and Axum refuses it. Scope the guard
in a block instead.
- `runtime_api::test_support::assert_runtime_stays_free` tests a handler which waits for a lock. Run
such a test with `#[tokio::test(flavor = "multi_thread", worker_threads = 1)]`.
### .NET App (`app/MindWork AI Studio/`)
**Entry point:** `app/MindWork AI Studio/Program.cs`
Key structure:
- **Program.cs** - Bootstraps Blazor Server, configures Kestrel, initializes encryption and Rust service
- **Provider/** - LLM provider implementations (OpenAI, Anthropic, Google, Mistral, etc.)
- `BaseProvider.cs` - Abstract base for all providers with streaming support
- `IProvider.cs` - Provider interface defining capabilities and streaming methods
- **Chat/** - Chat functionality and message handling
- **Assistants/** - Pre-configured assistants (translation, summarization, coding, etc.)
- `AssistantBase.razor` - Base component for all assistants
- **Agents/** - contains all agents, e.g., for data source selection, context validation, etc.
- `AgentDataSourceSelection.cs` - Selects appropriate data sources for queries
- `AgentRetrievalContextValidation.cs` - Validates retrieved context relevance
- **Tools/PluginSystem/** - Lua-based plugin system
- **Tools/Services/** - Core background services (settings, message bus, data sources, updates)
- **Tools/Rust/** - .NET wrapper for Rust API calls
- **Settings/** - Application settings and data models
- **Components/** - Reusable Blazor components
- **Pages/** - Top-level page components
### IPC Communication Flow
1. Rust runtime starts and generates TLS certificate
2. Rust starts internal HTTPS API on random port
3. Rust launches .NET sidecar, passing: API port, certificate fingerprint, API token, secret key
4. .NET reads environment variables and establishes secure HTTPS connection to Rust
5. .NET requests an app port from Rust, starts Blazor Server on that port
6. Rust opens Tauri webview pointing to localhost:app_port
7. Bi-directional communication: .NET ↔ Rust via HTTPS API
### Configuration and Metadata
## Configuration and Metadata
- `metadata.txt` - Build metadata (version, build time, component versions) read by both Rust and .NET
- `startup.env` - Development environment variables (generated by build script)
- `.NET project` reads metadata.txt at build time and injects as assembly attributes
## Plugin System
**Location:** `app/MindWork AI Studio/Plugins/`
Plugins are written in Lua and provide:
- **Language plugins** - I18N translations (e.g., German language pack)
- **Configuration plugins** - Enterprise IT configurations for centrally managed providers, settings
- **Assistant plugins** - custom assistants and direct-chat launchers, subject to approval or a local security audit
- **Model plugins** - what an organization's own models can do, see `documentation/Models.md`
**Example configuration plugin:** `app/MindWork AI Studio/Plugins/configuration/plugin.lua`
Plugins can configure:
- Self-hosted LLM providers
- Update behavior
- Preview features visibility
- Preselected profiles
- Chat templates
- etc.
Configuration plugins provide three kinds of values:
- **Managed settings:** simple values such as booleans, numbers, strings, enums, lists, or sets handled through `ManagedConfiguration`. These values may be locked or used as organization defaults. Which configuration plugin owns a locked setting is persisted in `Data.ManagedLockedConfigurations`, and organization defaults are tracked in `Data.ManagedEditableDefaults`. Both are cleaned up generically by `ManagedConfiguration.CleanupLeftOverManagedConfigurations(...)` when the owning plugin is gone. The value a setting had before a configuration plugin took it over is kept in `Data.ManagedUserValueSnapshots` and restored by that same clean-up, so removing a plugin hands the user's own value back instead of the app default.
- **Managed configuration objects:** complex Lua tables that are persisted into `SettingsManager.ConfigurationData`, implement `IConfigurationObject`, and are cleaned up through `PluginConfigurationObject.CleanLeftOverConfigurationObjects(...)`. Examples include providers, profiles, chat templates, data sources, and document analysis policies.
- **Live plugin content:** complex Lua tables that implement `ILivePluginContent` and are read live from running plugins instead of being persisted to `ConfigurationData`. Examples include `MANDATORY_INFOS` and `INTRODUCTIONS`. If live plugin content creates persistent side data, add a dedicated cleanup path for that side data, like mandatory-info acceptances.
When adding configuration plugin capabilities:
- For managed settings, update the corresponding data class in `app/MindWork AI Studio/Settings/DataModel/` to call `ManagedConfiguration.Register(...)` and process the setting in `PluginConfiguration.TryProcessConfiguration`. Cleaning up the setting when its configuration plugin was removed needs no extra step: `ManagedConfiguration.CleanupLeftOverManagedConfigurations(...)` iterates all registered settings. Do not add per-setting cleanup calls to `PluginFactory.Loading.LoadAll`.
- For managed configuration objects, update `PluginConfigurationObject.cs` and `PluginConfigurationObjectType.cs`, persist them in the appropriate `ConfigurationData` collection, and add cleanup via `PluginConfigurationObject.CleanLeftOverConfigurationObjects(...)`.
- For live plugin content, add a data type implementing `ILivePluginContent`, parse it in `PluginConfiguration`, expose it through `PluginFactory`, and add any required cleanup only for persistent side data.
- Always document the new capability in `app/MindWork AI Studio/Plugins/configuration/plugin.lua`.
## Tool Calling System
**Documentation:** `documentation/Tools.md`
When adding, changing, or removing model-driven tools, keep these parts in sync:
- `app/MindWork AI Studio/Tools/ToolCallingSystem/ToolCallingImplementations/` for the `IToolImplementation` class, which states its own `ToolDefinition` through `GetDefinition()`, written with `ToolSettingsSchemaBuilder` for its settings and `ToolParameterSchemaBuilder` for the arguments the model passes. There are no tool definition files; a tool arriving from elsewhere brings an `IToolDefinitionSource` instead.
- `app/MindWork AI Studio/Program.cs` for DI registration of the implementation. Registering it as an `IToolImplementation` is enough, because `CodeToolDefinitionSource` collects the definitions of all of them.
- An `IToolCollection` next to the tools, registered in `Program.cs`, when tools only make sense together. Selections, `DataTools.DisabledToolIds`, and the minimum provider confidence name tool collections; a tool outside a declared collection forms one under its own ID, and the ID of a tool inside one stands for its whole collection. Normalize a selection with `ToolRegistry.NormalizeSelection`, and expand it into tools with `ToolRegistry.ExpandSelection` only where the view of the model counts.
- `app/MindWork AI Studio/Tools/ToolCallingSystem/ToolSelectionRules.cs` when the shared tool-call limits change. A tool's own minimum provider confidence belongs in its definition, or in the definition of its collection, not here.
- `app/MindWork AI Studio/Tools/ToolCallingSystem/ToolSettingsOptionSources.cs` when a tool setting offers a fixed choice the app maintains, such as languages. Prefer this over spelling the values out in the settings schema; it keeps the list in one place and gives the user translated names.
- `app/MindWork AI Studio/Plugins/configuration/plugin.lua` to document each setting's field name, meaning, and data type. Tool settings need no code to be centrally manageable: an organization addresses them by `"<toolId>.<fieldName>"` in `DataTools.LockedToolSettings` or `DataTools.DefaultToolSettings`.
Tool implementations must treat model-provided arguments as untrusted input. Validate settings and arguments, protect secrets with `SensitiveTraceArgumentNames`, use `ToolExecutionBlockedException` for intentional policy blocks, and check provider confidence before returning sensitive data to the model.
A tool which belongs to a preview feature returns false from `IToolImplementation.IsAvailable` while the preview is switched off; the registry then leaves it out of every list, every request, and the token count, so no component has to check that preview for the tool. Every tool declares in `IToolImplementation.OutboundData` where its arguments go: a chat which read from a mailbox keeps the tools whose data goes further than the mailbox allows from being offered and from running, see `ToolSelectionRules.IsOutboundDataAllowed`. A tool which brings content of a mailbox into the chat raises `ToolExecutionResult.RequiredOutboundDataRestriction`, next to `RequiredProviderConfidence` and `RequiredDataSecurity`. "Searching Mailboxes" in `documentation/Tools.md` explains the mail tools.
A tool which offers itself from the context of a chat instead of being selected, such as `semantic_search`, sets `Activation = ToolActivation.CONTEXT` and tailors its function to each request in `ResolveFunctionAsync`. Code which decides something on behalf of a request — whether the classic RAG process steps back, say — asks `ToolRegistry.GetOfferBlockReasonAsync` or `ToolRegistry.GetEffectiveRetrievalModeAsync` with the provider settings of the request (`IProvider.CreateSettingsProvider`), never a check of its own: two answers which drift apart leave a chat searching nothing or twice.
## Model Capabilities
**Documentation:** `documentation/Models.md`
What a model can do is answered in `app/MindWork AI Studio/Models/`, through `provider.GetModelProfile()`. Never ask `ModelRegistry` directly from a component: the extension method is what adds the expert settings and what a provider's model list reported, and the registry alone answers neither.
When adding, changing, or removing model knowledge, keep these parts in sync:
- `app/MindWork AI Studio/Models/<Vendor>/<Family>.cs` for the family itself. Creating the class is enough — the source generator in `app/SourceGeneratedMappings/` collects every non-abstract `ModelFamily` and `IModelHost` at compile time, so there is no registration list. Do not add reflection here; `PublishTrimmed` is on.
- `app/Tests/Models/Corpus/` for the model IDs the family covers, marked as either unchanged or expected to change. A porting difference which nobody declared is what the corpus exists to catch.
- `app/MindWork AI Studio/Models/Kinds/` when the change is about what kind of model something is, rather than what it can do. These are ordinary rules of the same engine.
- `app/MindWork AI Studio/Models/Hosting/Hosts/` when a provider wraps model names or cannot pass an API through. A host unwraps and trims the transport; it states nothing about the model itself.
- `app/MindWork AI Studio/Plugins/models/plugin.lua` when a new field can be declared by an organization, and `app/MindWork AI Studio/Plugins/configuration/plugin.lua` when it can be overridden per provider instance.
Rules are never tried in order: specificity is computed from the rule, and two rules of equal specificity on one name fail the test suite. State how a model reasons with `Reasoning(...)` — the three reasoning capabilities are override vocabulary and must never appear in a profile. Every family and every host has to name the page it was read from and the day somebody read it; `dotnet run verify-models` reports the ones which have gone stale.
## RAG (Retrieval-Augmented Generation)
RAG is available as a beta preview feature. Architecture:
- **External Retrieval Interface (ERI)** - Contract for integrating external data sources
- **Data Sources** - Local files and external data via ERI servers
- **Two ways to search** - By default, the chat model searches the data sources itself through the tool `semantic_search`, whenever a question calls for it. The classic process (`AISrcSelWithRetCtxVal`) searches them with every message instead, when the user chose so per chat (`DataSourceOptions.RetrievalMode`) or whenever the tool cannot be offered. `ToolRegistry.GetEffectiveRetrievalModeAsync` decides between the two; pass its answer to `DataSourceService`, because the agents only count as providers that see the data when they actually run. See "Searching Data Sources" in `documentation/Tools.md`.
- **Agents** - AI agents select data sources and validate retrieval quality, in the classic process only
- **Embedding providers** - Support for various embedding models
- **Vector database** - Qdrant Edge, embedded in the Rust runtime; see "Databases" below
- **Index database** - SQLite, holding the file fingerprints and the chunk texts for full-text search; see "Databases" below
- **File processing** - Extracts text from PDF, DOCX, XLSX via Rust runtime
### Indexed data sources
Everything AI Studio embeds itself runs through one pipeline in `app/MindWork AI Studio/Tools/Services/Indexing/`,
driven by `DataSourceEmbeddingService`, which queues the runs, prepares each one and owns the statuses. A new
kind of data source plugs into this pipeline instead of building its own.
The parts:
- **`IIndexedDataSource`** (`Settings/`) - what indexing needs to know about a data source: confidence level,
embedding provider and chunk settings. `IDataSourceBase` is what every data source has. Implement
`IDataSource` on top only when classic RAG, Semantic Search and the agents should see the data source. A data
source kept in a list of its own implements `IIndexedDataSource` alone, and the compiler keeps it out of
`DataSources`.
- **`IIndexedSourceIndexer`** - one per kind of data source. `Supports` claims the data sources, `ProcessAsync`
finds and reads the documents of one run, and `TrackChanges` / `StopTracking` notice changes on their own.
`FileSourceIndexer` is the reference: a file system watcher per data source, fingerprints over path, size and
write time.
- **`IndexedRunContext`** - one prepared run: both stores, the embedding provider, the manifest and the
collection. `IndexDocumentAsync` embeds and stores one document; the cleanup methods remove what a failed
attempt left behind.
- **`EmbeddingDocument`** - one document: its key, its index row, its display name, and how to read its chunks.
- **`DocumentRunProgress`** - counts the documents, records indexed and failed ones in the stores, publishes the
status and completes the run.
- **`TextChunker`** - cuts text into chunks the embedding provider accepts. Pick one of its strategies; do not
write a chunker of your own.
To add a kind of data source:
1. Write its indexer in `Tools/Services/Indexing/` and create it in `DataSourceEmbeddingService.CreateIndexers`,
which hands every indexer the same `TextChunker`.
2. Gate it in `IsSupportedIndexedSource`, behind a preview feature of its own while it is new.
3. When the data source is not kept in `DataSources`, add its list to `GetConfiguredIndexedSources`. Every lookup
by id and every pass over all data sources goes through it: the startup hash check,
`QueueAllInternalDataSourcesAsync` and `RefreshWatchers`.
4. Keep whatever the kind has to remember beyond its documents in tables of its own in the index store, added
by an EF Core migration (see "Databases").
5. Report every status through `DocumentRunProgress`, so all rows of the embedding page behave alike.
Rules which are easy to break:
- **A document key is not a path.** Only files use their full path as the key. Never pass a key through the
`Path` APIs: on Windows, `Path.GetFullPath` reads a key like `mail:…` as a file with an alternate data stream.
- **Ids and signature are pinned.** The formats in `IndexedDocumentIds` and the embedding signature
(`DataSourceEmbeddingService.BuildEmbeddingSignature`) are fixed by tests, because every stored chunk and every
index depends on them. When the metadata stored next to a chunk changes, raise `CHUNK_METADATA_VERSION`
deliberately: that rebuilds every index.
- **Every content path goes through the prompt injection filter.** Files pass the sanitizer of the runtime while
their text is extracted; a new kind of data source needs its own pass through `PromptInjectionGuardService`.
- **The service decides when, the indexer decides how.** Whether changes are tracked at all depends on the
automatic refresh setting and the startup hash check, and only the embedding service decides that.
- **Moving a `TB()` text into another class gives it a new I18N key**, so its translation is made anew during
the next localization run.
## Databases
Local RAG runs on two databases, addressed through `DatabaseRole`:
- **`VECTOR_STORE`** — Qdrant Edge through the `qdrant-edge` crate, running **in-process inside the
Rust runtime**. There is no sidecar process, no port 6333 and no Qdrant API key; .NET reaches it
over the internal runtime API (`/system/qdrant-edge/*`, see `runtime/src/qdrant_edge_database.rs`),
secured by the same TLS and API token as every other runtime call. One store per data source,
named `rag_<data source guid>`, holding a single named vector `embedding` per point.
- **`INDEX_STORE`** — SQLite at `<data directory>/databases/sqlite/rag-index.sqlite3`, reached
through EF Core. It holds the data sources, the file fingerprints, the chunk texts and an FTS5
index over them, plus the files which permanently failed to index.
`DatabaseClientProvider` is the only way to a client. It caches one per role and guards each role
with its own semaphore, so never construct a client yourself.
When working on these, keep in mind:
- **`GetDisplayInfo()` feeds the information page.** A new diagnostic value belongs in the client
that knows it, not in `Pages/Information.razor.cs`. The page renders whatever label-value pairs it
receives and stays free of per-database knowledge.
- **Let every probe in `GetDisplayInfo()` catch its own failure.** When the method throws, the page
replaces the *entire* block with the fallback client, so one unreadable value costs all the others
as well.
- **Raw SQL against SQLite goes through `context.Database.GetDbConnection()`**, not through
`SqlQueryRaw<T>`: that one expects a column named `Value` and wraps the statement, so a `PRAGMA`
never works with it.
- **A new EF Core migration needs a `[DynamicDependency]`** in `IndexStoreSchemaMigrator`, because
`PublishTrimmed` is on and the migration type would otherwise be trimmed away. The "Schema version"
line on the information page shows the applied and pending counts, so a forgotten entry becomes
visible there.
- **Counts in the UI go through `long.CompactCount()` / `int.CompactCount()`** (`Tools/LongExtensions.cs`),
which shortens anything above 999 to `1.46k` or `4.51M` and formats it with the culture of the
active language plugin. Storage sizes are the exception: they keep using the byte formatters.
## Enterprise IT Support
AI Studio supports centralized configuration for enterprise environments:
- **Registry (Windows)** or **environment variables** (all platforms) specify configuration server URL and ID
- Configuration downloaded as ZIP containing Lua plugin
- Checks for updates every ~16 minutes via ETag
- Allows IT departments to pre-configure providers, settings, and chat templates
**Documentation:** `documentation/Enterprise IT.md`
## Provider Confidence System
Multi-level confidence scheme allows users to control which providers see which data:
- Confidence levels: e.g. `NONE`, `LOW`, `MEDIUM`, `HIGH`, and some more granular levels
- Each assistant/feature can require a minimum confidence level
- Users assign confidence levels to providers based on trust
**Implementation:** `app/MindWork AI Studio/Provider/Confidence.cs`
## Dependencies and Frameworks
**Rust:**
- Tauri 2 - Desktop application framework
- Axum - HTTPS API server
- tokio - Async runtime
- keyring - OS keyring integration
- pdfium-render - PDF text extraction
- calamine - Excel file parsing
- qdrant-edge - Embedded vector database
**.NET:**
- Blazor Server - UI framework
- MudBlazor - Component library
- LuaCSharp - Lua scripting engine
- HtmlAgilityPack - HTML parsing
- ReverseMarkdown - HTML to Markdown conversion
- EF Core Sqlite + SQLitePCLRaw - the local RAG index
## Security
- **Encryption:** AES-256-CBC with PBKDF2 key derivation for sensitive data
- **IPC:** TLS-secured communication with random ports and API tokens
- **Secrets:** OS keyring for persistent secret storage (API keys, etc.)
- **Sandboxing:** Tauri provides OS-level sandboxing
## Release Process
1. Create changelog file: `app/MindWork AI Studio/wwwroot/changelog/vX.Y.Z.md`
2. Check that every contribution in the release is credited, see "Crediting contributors" below. The
authors of all pull requests merged since the last release are listed by
`gh pr list -R MindWorkAI/AI-Studio --state merged --limit 1000 --search "merged:>=<YYYY-MM-DD>" --json author --jq '[.[].author.login] | unique | .[]'`,
with the date of the last release tag from `git log -1 --format=%cs <last release tag>`; skip bots
3. Commit changelog
4. Run from `app/Build`: `dotnet run release --action <build|month|year>`
5. Create PR with version bump and changes
6. After PR merge, maintainer creates git tag: `vX.Y.Z`
7. GitHub Actions builds release binaries for all platforms
8. Binaries uploaded to GitHub Releases
## Localization
The app's texts are localized in two steps, and the developer always does the first one.
1. The developer starts the app, which runs the I18N collector, and runs the localization assistant
in the app for German and US English. Agents never write these initial translations themselves:
they neither add nor regenerate entries in `app/MindWork AI Studio/Assistants/I18N/allTexts.lua`,
`app/MindWork AI Studio/Plugins/languages/en-us-97dfb1ba-50c4-4440-8dfa-6575daf543c8/plugin.lua`,
or `app/MindWork AI Studio/Plugins/languages/de-de-43065dbc-78d0-45b7-92be-f14c2926e2dc/plugin.lua`.
When new or changed texts are waiting for translation, remind the developer to start the app and
run the localization.
2. Afterward, agents always review the German translation. Compare the new and changed values of the
de-de `plugin.lua` with `main`, check them against the wording already established there, and
correct or improve them directly in that file. `allTexts.lua` and the en-us `plugin.lua` stay as
the assistant wrote them.
## Important Development Notes
- **File changes require Write/Edit tools** - Never use bash commands like `cat <<EOF` or `echo >`
- **End of file formatting** - Do not append an extra empty line at the end of files.
- **No automated formatting for Rust or .NET files** - Never run automated formatters on Rust files (`.rs`) or .NET files (`.cs`, `.razor`, `.csproj`, etc.). Only make the minimal manual formatting changes required for the specific edit.
- **I18N resources are generated** - The developer produces the translations by running the localization assistant in the app; agents only review and correct the German values afterward. See "Localization" above.
- **Spaces in paths** - Always quote paths with spaces in bash commands
- **Agent-run builds** - Never start `.NET` or Rust builds in the agent's own shell; it is sandboxed. Use the `rider` and `rustrover` MCP servers instead, which build in the IDE outside that sandbox. See "Running builds from an agent" above.
- **Debug environment** - Reads `startup.env` file with IPC credentials
- **Production environment** - Runtime launches .NET sidecar with environment variables
- **MudBlazor** - Component library requires DI setup in Program.cs
- **Encryption** - Initialized before Rust service is marked ready
- **Message Bus** - Singleton event bus for cross-component communication inside the .NET app
- **Naming conventions** - Constants, enum members, and `static readonly` fields use `UPPER_SNAKE_CASE` such as `MY_CONSTANT`.
- **Compatibility shims** - Temporary fallback or read-repair code must be documented in `documentation/compatibility-shims/` with an introduced date, remove-after date, code references, and removal checklist. Add a short code comment near the shim that references the document and remove-after date. Check this folder before adding similar fallback logic, and do not extend expired shims without explicit maintainer direction. Do not use this process for permanent settings schema migrations; those belong in `app/MindWork AI Studio/Settings/SettingsMigrations.cs`.
- **Empty lines** - Avoid adding extra empty lines at the end of files.
## Changelogs
Changelogs are located in `app/MindWork AI Studio/wwwroot/changelog/` with filenames `vX.Y.Z.md`. These changelogs are meant to be for normal end-users
@@ -473,34 +173,6 @@ inside the entry instead, even when that repeats a few words from another one.
**Split a topic into several short entries** rather than growing a single long one, and address the
reader with "you".
### Crediting contributors
When a pull request is merged, or a release is prepared, check both places where we thank contributors.
This holds for everybody except the maintainer, members of the core team included, not only for external
contributors:
- **The changelog entry of the change.** Thank the contributor at the end of the entry, in the form
``<first name> <last name> (`<GitHub username>`)``, and call a first contribution out as such.
- **The "Code Contributions" list on the supporters page** in `app/MindWork AI Studio/Pages/Supporters.razor`.
Add contributors who are not listed yet, one
`<Supporter Name="<GitHub username>" Type="SupporterType.INDIVIDUAL" URL="https://github.com/<GitHub username>" Acknowledgment="@T("…")"/>`
each, at the end of the list. The list is sorted by the first merged contribution of each person; someone
whom the changelog thanks for work inside another person's pull request counts from that pull request on.
The acknowledgment calls the person by their first name, unless the credit choice below rules out their
name, and says what they built. It never compares contributors or counts their contributions, and the
order of the list is chronological only. Bug reports
alone and commissioned work outside this repository get no entry on the supporters page; the changelog may
still thank them.
The credit choice in the pull request template is binding:
- **"Credit me with my GitHub username only":** use the GitHub username alone. No real name, neither in
the changelog nor in the acknowledgment text.
- **"Do not credit me":** neither a changelog mention nor an entry on the supporters page.
- **No choice ticked, or a pull request from before the template:** the GitHub username plus the name, when
the contributor shows it publicly on their GitHub profile (`gh api users/<GitHub username> --jq .name`)
or an earlier changelog already thanked them by that name.
Acknowledgments on the supporters page are `T()` texts, so the two steps of "Localization" above apply:
remind the developer to run the localization, then review the German value.
**Thanking contributors.** Before an entry thanks a contributor, and whenever a pull request is merged,
read "Crediting contributors" in `app/Build/AGENTS.md` first: the credit choice in the pull request
template decides whether and how we name somebody.
+151
View File
@@ -0,0 +1,151 @@
# AGENTS.md
These instructions cover the .NET app: all C#, Razor, and Lua code under `app/`, and everything else
there. They add to the `AGENTS.md` in the repository root, which applies here as well. Several areas below
`app/` have guides of their own; the table "Area guides" in the root `AGENTS.md` lists them.
## .NET App (`app/MindWork AI Studio/`)
**Entry point:** `app/MindWork AI Studio/Program.cs`
Key structure:
- **Program.cs** - Bootstraps Blazor Server, configures Kestrel, initializes encryption and Rust service
- **Provider/** - LLM provider implementations (OpenAI, Anthropic, Google, Mistral, etc.)
- `BaseProvider.cs` - Abstract base for all providers with streaming support
- `IProvider.cs` - Provider interface defining capabilities and streaming methods
- **Chat/** - Chat functionality and message handling
- **Assistants/** - Pre-configured assistants (translation, summarization, coding, etc.)
- `AssistantBase.razor` - Base component for all assistants
- **Agents/** - contains all agents, e.g., for data source selection, context validation, etc.
- `AgentDataSourceSelection.cs` - Selects appropriate data sources for queries
- `AgentRetrievalContextValidation.cs` - Validates retrieved context relevance
- **Tools/PluginSystem/** - Lua-based plugin system
- **Tools/Services/** - Core background services (settings, message bus, data sources, updates)
- **Tools/Rust/** - .NET wrapper for Rust API calls
- **Settings/** - Application settings and data models
- **Components/** - Reusable Blazor components
- **Pages/** - Top-level page components
## Important Development Notes
- **Naming conventions** - Constants, enum members, and `static readonly` fields use `UPPER_SNAKE_CASE` such as `MY_CONSTANT`.
- **Message Bus** - Singleton event bus for cross-component communication inside the .NET app
- **Encryption** - Initialized before Rust service is marked ready
- **Debug environment** - Reads `startup.env` file with IPC credentials
- **Production environment** - Runtime launches .NET sidecar with environment variables
## Tests
An assembly-wide `[SetUpFixture]` in `app/Tests/TestHost.cs` fills the static application state that
the app itself only fills while starting up, `Program.LOGGER_FACTORY` above all. Types that
initialize a static logger from it — `Settings.Provider` among them — otherwise die in their type
initializer before the first assertion. Prefer writing new code so that it does not reach for such
statics at all.
## Localization
The app's texts are localized in two steps, and the developer always does the first one.
1. The developer starts the app, which runs the I18N collector, and runs the localization assistant
in the app for German and US English. Agents never write these initial translations themselves:
they neither add nor regenerate entries in `app/MindWork AI Studio/Assistants/I18N/allTexts.lua`,
`app/MindWork AI Studio/Plugins/languages/en-us-97dfb1ba-50c4-4440-8dfa-6575daf543c8/plugin.lua`,
or `app/MindWork AI Studio/Plugins/languages/de-de-43065dbc-78d0-45b7-92be-f14c2926e2dc/plugin.lua`.
When new or changed texts are waiting for translation, remind the developer to start the app and
run the localization.
2. Afterward, agents always review the German translation. Compare the new and changed values of the
de-de `plugin.lua` with `main`, check them against the wording already established there, and
correct or improve them directly in that file. `allTexts.lua` and the en-us `plugin.lua` stay as
the assistant wrote them.
**Moving a `TB()` text into another class gives it a new I18N key**, so its translation is made anew during
the next localization run.
## Plugin System
**Location:** `app/MindWork AI Studio/Plugins/`
Plugins are written in Lua and provide:
- **Language plugins** - I18N translations (e.g., German language pack)
- **Configuration plugins** - Enterprise IT configurations for centrally managed providers, settings
- **Assistant plugins** - custom assistants and direct-chat launchers, subject to approval or a local security audit
- **Model plugins** - what an organization's own models can do, see `documentation/Models.md`
**Example configuration plugin:** `app/MindWork AI Studio/Plugins/configuration/plugin.lua`
**Area guide:** `app/MindWork AI Studio/Tools/PluginSystem/AGENTS.md`, with the three kinds of values a
configuration plugin provides and how to add each of them.
## Tool Calling System
**Documentation:** `documentation/Tools.md`
Selections, `DataTools.DisabledToolIds`, and the minimum provider confidence name tool collections; a tool outside a declared collection forms one under its own ID, and the ID of a tool inside one stands for its whole collection. Normalize a selection with `ToolRegistry.NormalizeSelection`, and expand it into tools with `ToolRegistry.ExpandSelection` only where the view of the model counts.
Code which decides something on behalf of a request — whether the classic RAG process steps back, say — asks `ToolRegistry.GetOfferBlockReasonAsync` or `ToolRegistry.GetEffectiveRetrievalModeAsync` with the provider settings of the request (`IProvider.CreateSettingsProvider`), never a check of its own: two answers which drift apart leave a chat searching nothing or twice.
**Area guide:** `app/MindWork AI Studio/Tools/ToolCallingSystem/AGENTS.md`, with how to add, change, or
remove a tool and the rules every tool implementation follows.
## Model Capabilities
**Documentation:** `documentation/Models.md`
What a model can do is answered in `app/MindWork AI Studio/Models/`, through `provider.GetModelProfile()`. Never ask `ModelRegistry` directly from a component: the extension method is what adds the expert settings and what a provider's model list reported, and the registry alone answers neither.
**Area guide:** `app/MindWork AI Studio/Models/AGENTS.md`, with how to add, change, or remove model
knowledge.
## RAG (Retrieval-Augmented Generation)
RAG is available as a beta preview feature. Architecture:
- **External Retrieval Interface (ERI)** - Contract for integrating external data sources
- **Data Sources** - Local files and external data via ERI servers
- **Two ways to search** - By default, the chat model searches the data sources itself through the tool `semantic_search`, whenever a question calls for it. The classic process (`AISrcSelWithRetCtxVal`) searches them with every message instead, when the user chose so per chat (`DataSourceOptions.RetrievalMode`) or whenever the tool cannot be offered. `ToolRegistry.GetEffectiveRetrievalModeAsync` decides between the two; pass its answer to `DataSourceService`, because the agents only count as providers that see the data when they actually run. See "Searching Data Sources" in `documentation/Tools.md`.
- **Agents** - AI agents select data sources and validate retrieval quality, in the classic process only
- **Embedding providers** - Support for various embedding models
- **Vector database** - Qdrant Edge, embedded in the Rust runtime; see "Databases" below
- **Index database** - SQLite, holding the file fingerprints and the chunk texts for full-text search; see "Databases" below
- **File processing** - Extracts text from PDF, DOCX, XLSX via Rust runtime
### Indexed data sources
Everything AI Studio embeds itself runs through one pipeline in `app/MindWork AI Studio/Tools/Services/Indexing/`,
driven by `DataSourceEmbeddingService`, which queues the runs, prepares each one and owns the statuses. A new
kind of data source plugs into this pipeline instead of building its own.
**Every content path goes through the prompt injection filter.** Files pass the sanitizer of the runtime while
their text is extracted; a new kind of data source needs its own pass through `PromptInjectionGuardService`.
**Area guide:** `app/MindWork AI Studio/Tools/Services/Indexing/AGENTS.md`, with the parts of the
pipeline, how to add a kind of data source, and the rules which are easy to break.
## Databases
`DatabaseClientProvider` is the only way to a client. It caches one per role and guards each role
with its own semaphore, so never construct a client yourself.
**Counts in the UI go through `long.CompactCount()` / `int.CompactCount()`** (`Tools/LongExtensions.cs`),
which shortens anything above 999 to `1.46k` or `4.51M` and formats it with the culture of the
active language plugin. Storage sizes are the exception: they keep using the byte formatters.
**Area guide:** `app/MindWork AI Studio/Tools/Databases/AGENTS.md`, with both stores and what to keep in
mind when working on them, EF Core migrations included.
## Enterprise IT Support
AI Studio supports centralized configuration for enterprise environments:
- **Registry (Windows)** or **environment variables** (all platforms) specify configuration server URL and ID
- Configuration downloaded as ZIP containing Lua plugin
- Checks for updates every ~16 minutes via ETag
- Allows IT departments to pre-configure providers, settings, and chat templates
**Documentation:** `documentation/Enterprise IT.md`
## Provider Confidence System
Multi-level confidence scheme allows users to control which providers see which data:
- Confidence levels: e.g. `NONE`, `LOW`, `MEDIUM`, `HIGH`, and some more granular levels
- Each assistant/feature can require a minimum confidence level
- Users assign confidence levels to providers based on trust
**Implementation:** `app/MindWork AI Studio/Provider/Confidence.cs`
+51
View File
@@ -0,0 +1,51 @@
# AGENTS.md
These instructions cover preparing a release and crediting contributors. They add to the `AGENTS.md` in
the repository root, which applies here as well; its "Changelogs" section explains how to write the
changelog entries.
## Release Process
1. Create changelog file: `app/MindWork AI Studio/wwwroot/changelog/vX.Y.Z.md`
2. Check that every contribution in the release is credited, see "Crediting contributors" below. The
authors of all pull requests merged since the last release are listed by
`gh pr list -R MindWorkAI/AI-Studio --state merged --limit 1000 --search "merged:>=<YYYY-MM-DD>" --json author --jq '[.[].author.login] | unique | .[]'`,
with the date of the last release tag from `git log -1 --format=%cs <last release tag>`; skip bots
3. Commit changelog
4. Run from `app/Build`: `dotnet run release --action <build|month|year>`
5. Create PR with version bump and changes
6. After PR merge, maintainer creates git tag: `vX.Y.Z`
7. GitHub Actions builds release binaries for all platforms
8. Binaries uploaded to GitHub Releases
## Crediting contributors
When a pull request is merged, or a release is prepared, check both places where we thank contributors.
This holds for everybody except the maintainer, members of the core team included, not only for external
contributors:
- **The changelog entry of the change.** Thank the contributor at the end of the entry, in the form
``<first name> <last name> (`<GitHub username>`)``, and call a first contribution out as such.
- **The "Code Contributions" list on the supporters page** in `app/MindWork AI Studio/Pages/Supporters.razor`.
Add contributors who are not listed yet, one
`<Supporter Name="<GitHub username>" Type="SupporterType.INDIVIDUAL" URL="https://github.com/<GitHub username>" Acknowledgment="@T("…")"/>`
each, at the end of the list. The list is sorted by the first merged contribution of each person; someone
whom the changelog thanks for work inside another person's pull request counts from that pull request on.
The acknowledgment calls the person by their first name, unless the credit choice below rules out their
name, and says what they built. It never compares contributors or counts their contributions, and the
order of the list is chronological only. Bug reports
alone and commissioned work outside this repository get no entry on the supporters page; the changelog may
still thank them.
The credit choice in the pull request template is binding:
- **"Credit me with my GitHub username only":** use the GitHub username alone. No real name, neither in
the changelog nor in the acknowledgment text.
- **"Do not credit me":** neither a changelog mention nor an entry on the supporters page.
- **No choice ticked, or a pull request from before the template:** the GitHub username plus the name, when
the contributor shows it publicly on their GitHub profile (`gh api users/<GitHub username> --jq .name`)
or an earlier changelog already thanked them by that name.
Acknowledgments on the supporters page are `T()` texts, so the two steps of "Localization" in the root
`AGENTS.md` apply: remind the developer to run the localization, then review the German value.
+1
View File
@@ -0,0 +1 @@
@AGENTS.md
+1
View File
@@ -0,0 +1 @@
@AGENTS.md
+15
View File
@@ -0,0 +1,15 @@
# AGENTS.md
These instructions cover the model knowledge in `app/MindWork AI Studio/Models/`. They add to the
`AGENTS.md` in the repository root, which applies here as well.
## Model Capabilities
When adding, changing, or removing model knowledge, keep these parts in sync:
- `app/MindWork AI Studio/Models/<Vendor>/<Family>.cs` for the family itself. Creating the class is enough — the source generator in `app/SourceGeneratedMappings/` collects every non-abstract `ModelFamily` and `IModelHost` at compile time, so there is no registration list. Do not add reflection here; `PublishTrimmed` is on.
- `app/Tests/Models/Corpus/` for the model IDs the family covers, marked as either unchanged or expected to change. A porting difference which nobody declared is what the corpus exists to catch.
- `app/MindWork AI Studio/Models/Kinds/` when the change is about what kind of model something is, rather than what it can do. These are ordinary rules of the same engine.
- `app/MindWork AI Studio/Models/Hosting/Hosts/` when a provider wraps model names or cannot pass an API through. A host unwraps and trims the transport; it states nothing about the model itself.
- `app/MindWork AI Studio/Plugins/models/plugin.lua` when a new field can be declared by an organization, and `app/MindWork AI Studio/Plugins/configuration/plugin.lua` when it can be overridden per provider instance.
Rules are never tried in order: specificity is computed from the rule, and two rules of equal specificity on one name fail the test suite. State how a model reasons with `Reasoning(...)` — the three reasoning capabilities are override vocabulary and must never appear in a profile. Every family and every host has to name the page it was read from and the day somebody read it; `dotnet run verify-models` reports the ones which have gone stale.
+1
View File
@@ -0,0 +1 @@
@AGENTS.md
@@ -0,0 +1,33 @@
# AGENTS.md
These instructions cover the databases in `app/MindWork AI Studio/Tools/Databases/`. They add to the
`AGENTS.md` in the repository root, which applies here as well.
## Databases
Local RAG runs on two databases, addressed through `DatabaseRole`:
- **`VECTOR_STORE`** — Qdrant Edge through the `qdrant-edge` crate, running **in-process inside the
Rust runtime**. There is no sidecar process, no port 6333 and no Qdrant API key; .NET reaches it
over the internal runtime API (`/system/qdrant-edge/*`, see `runtime/src/qdrant_edge_database.rs`),
secured by the same TLS and API token as every other runtime call. One store per data source,
named `rag_<data source guid>`, holding a single named vector `embedding` per point.
- **`INDEX_STORE`** — SQLite at `<data directory>/databases/sqlite/rag-index.sqlite3`, reached
through EF Core. It holds the data sources, the file fingerprints, the chunk texts and an FTS5
index over them, plus the files which permanently failed to index.
When working on these, keep in mind:
- **`GetDisplayInfo()` feeds the information page.** A new diagnostic value belongs in the client
that knows it, not in `Pages/Information.razor.cs`. The page renders whatever label-value pairs it
receives and stays free of per-database knowledge.
- **Let every probe in `GetDisplayInfo()` catch its own failure.** When the method throws, the page
replaces the *entire* block with the fallback client, so one unreadable value costs all the others
as well.
- **Raw SQL against SQLite goes through `context.Database.GetDbConnection()`**, not through
`SqlQueryRaw<T>`: that one expects a column named `Value` and wraps the statement, so a `PRAGMA`
never works with it.
- **A new EF Core migration needs a `[DynamicDependency]`** in `IndexStoreSchemaMigrator`, because
`PublishTrimmed` is on and the migration type would otherwise be trimmed away. The "Schema version"
line on the information page shows the applied and pending counts, so a forgotten entry becomes
visible there.
@@ -0,0 +1 @@
@AGENTS.md
@@ -0,0 +1,26 @@
# AGENTS.md
These instructions cover the plugin system in `app/MindWork AI Studio/Tools/PluginSystem/` and what
configuration plugins can provide. They add to the `AGENTS.md` in the repository root, which applies here
as well.
## Configuration plugins
Plugins can configure:
- Self-hosted LLM providers
- Update behavior
- Preview features visibility
- Preselected profiles
- Chat templates
- etc.
Configuration plugins provide three kinds of values:
- **Managed settings:** simple values such as booleans, numbers, strings, enums, lists, or sets handled through `ManagedConfiguration`. These values may be locked or used as organization defaults. Which configuration plugin owns a locked setting is persisted in `Data.ManagedLockedConfigurations`, and organization defaults are tracked in `Data.ManagedEditableDefaults`. Both are cleaned up generically by `ManagedConfiguration.CleanupLeftOverManagedConfigurations(...)` when the owning plugin is gone. The value a setting had before a configuration plugin took it over is kept in `Data.ManagedUserValueSnapshots` and restored by that same clean-up, so removing a plugin hands the user's own value back instead of the app default.
- **Managed configuration objects:** complex Lua tables that are persisted into `SettingsManager.ConfigurationData`, implement `IConfigurationObject`, and are cleaned up through `PluginConfigurationObject.CleanLeftOverConfigurationObjects(...)`. Examples include providers, profiles, chat templates, data sources, and document analysis policies.
- **Live plugin content:** complex Lua tables that implement `ILivePluginContent` and are read live from running plugins instead of being persisted to `ConfigurationData`. Examples include `MANDATORY_INFOS` and `INTRODUCTIONS`. If live plugin content creates persistent side data, add a dedicated cleanup path for that side data, like mandatory-info acceptances.
When adding configuration plugin capabilities:
- For managed settings, update the corresponding data class in `app/MindWork AI Studio/Settings/DataModel/` to call `ManagedConfiguration.Register(...)` and process the setting in `PluginConfiguration.TryProcessConfiguration`. Cleaning up the setting when its configuration plugin was removed needs no extra step: `ManagedConfiguration.CleanupLeftOverManagedConfigurations(...)` iterates all registered settings. Do not add per-setting cleanup calls to `PluginFactory.Loading.LoadAll`.
- For managed configuration objects, update `PluginConfigurationObject.cs` and `PluginConfigurationObjectType.cs`, persist them in the appropriate `ConfigurationData` collection, and add cleanup via `PluginConfigurationObject.CleanLeftOverConfigurationObjects(...)`.
- For live plugin content, add a data type implementing `ILivePluginContent`, parse it in `PluginConfiguration`, expose it through `PluginFactory`, and add any required cleanup only for persistent side data.
- Always document the new capability in `app/MindWork AI Studio/Plugins/configuration/plugin.lua`.
@@ -0,0 +1 @@
@AGENTS.md
@@ -0,0 +1,47 @@
# AGENTS.md
These instructions cover the indexing pipeline in `app/MindWork AI Studio/Tools/Services/Indexing/` and
`DataSourceEmbeddingService`. They add to the `AGENTS.md` in the repository root, which applies here as
well.
## Indexed data sources
The parts:
- **`IIndexedDataSource`** (`Settings/`) - what indexing needs to know about a data source: confidence level,
embedding provider and chunk settings. `IDataSourceBase` is what every data source has. Implement
`IDataSource` on top only when classic RAG, Semantic Search and the agents should see the data source. A data
source kept in a list of its own implements `IIndexedDataSource` alone, and the compiler keeps it out of
`DataSources`.
- **`IIndexedSourceIndexer`** - one per kind of data source. `Supports` claims the data sources, `ProcessAsync`
finds and reads the documents of one run, and `TrackChanges` / `StopTracking` notice changes on their own.
`FileSourceIndexer` is the reference: a file system watcher per data source, fingerprints over path, size and
write time.
- **`IndexedRunContext`** - one prepared run: both stores, the embedding provider, the manifest and the
collection. `IndexDocumentAsync` embeds and stores one document; the cleanup methods remove what a failed
attempt left behind.
- **`EmbeddingDocument`** - one document: its key, its index row, its display name, and how to read its chunks.
- **`DocumentRunProgress`** - counts the documents, records indexed and failed ones in the stores, publishes the
status and completes the run.
- **`TextChunker`** - cuts text into chunks the embedding provider accepts. Pick one of its strategies; do not
write a chunker of your own.
To add a kind of data source:
1. Write its indexer in `Tools/Services/Indexing/` and create it in `DataSourceEmbeddingService.CreateIndexers`,
which hands every indexer the same `TextChunker`.
2. Gate it in `IsSupportedIndexedSource`, behind a preview feature of its own while it is new.
3. When the data source is not kept in `DataSources`, add its list to `GetConfiguredIndexedSources`. Every lookup
by id and every pass over all data sources goes through it: the startup hash check,
`QueueAllInternalDataSourcesAsync` and `RefreshWatchers`.
4. Keep whatever the kind has to remember beyond its documents in tables of its own in the index store, added
by an EF Core migration (see `app/MindWork AI Studio/Tools/Databases/AGENTS.md`).
5. Report every status through `DocumentRunProgress`, so all rows of the embedding page behave alike.
Rules which are easy to break:
- **A document key is not a path.** Only files use their full path as the key. Never pass a key through the
`Path` APIs: on Windows, `Path.GetFullPath` reads a key like `mail:…` as a file with an alternate data stream.
- **Ids and signature are pinned.** The formats in `IndexedDocumentIds` and the embedding signature
(`DataSourceEmbeddingService.BuildEmbeddingSignature`) are fixed by tests, because every stored chunk and every
index depends on them. When the metadata stored next to a chunk changes, raise `CHUNK_METADATA_VERSION`
deliberately: that rebuilds every index.
- **The service decides when, the indexer decides how.** Whether changes are tracked at all depends on the
automatic refresh setting and the startup hash check, and only the embedding service decides that.
@@ -0,0 +1 @@
@AGENTS.md
@@ -0,0 +1,21 @@
# AGENTS.md
These instructions cover the tool calling system in `app/MindWork AI Studio/Tools/ToolCallingSystem/`. They
add to the `AGENTS.md` in the repository root, which applies here as well; its "Tool Calling System"
section explains how selections name tool collections.
## Tool Calling System
When adding, changing, or removing model-driven tools, keep these parts in sync:
- `app/MindWork AI Studio/Tools/ToolCallingSystem/ToolCallingImplementations/` for the `IToolImplementation` class, which states its own `ToolDefinition` through `GetDefinition()`, written with `ToolSettingsSchemaBuilder` for its settings and `ToolParameterSchemaBuilder` for the arguments the model passes. There are no tool definition files; a tool arriving from elsewhere brings an `IToolDefinitionSource` instead.
- `app/MindWork AI Studio/Program.cs` for DI registration of the implementation. Registering it as an `IToolImplementation` is enough, because `CodeToolDefinitionSource` collects the definitions of all of them.
- An `IToolCollection` next to the tools, registered in `Program.cs`, when tools only make sense together.
- `app/MindWork AI Studio/Tools/ToolCallingSystem/ToolSelectionRules.cs` when the shared tool-call limits change. A tool's own minimum provider confidence belongs in its definition, or in the definition of its collection, not here.
- `app/MindWork AI Studio/Tools/ToolCallingSystem/ToolSettingsOptionSources.cs` when a tool setting offers a fixed choice the app maintains, such as languages. Prefer this over spelling the values out in the settings schema; it keeps the list in one place and gives the user translated names.
- `app/MindWork AI Studio/Plugins/configuration/plugin.lua` to document each setting's field name, meaning, and data type. Tool settings need no code to be centrally manageable: an organization addresses them by `"<toolId>.<fieldName>"` in `DataTools.LockedToolSettings` or `DataTools.DefaultToolSettings`.
Tool implementations must treat model-provided arguments as untrusted input. Validate settings and arguments, protect secrets with `SensitiveTraceArgumentNames`, use `ToolExecutionBlockedException` for intentional policy blocks, and check provider confidence before returning sensitive data to the model.
A tool which belongs to a preview feature returns false from `IToolImplementation.IsAvailable` while the preview is switched off; the registry then leaves it out of every list, every request, and the token count, so no component has to check that preview for the tool. Every tool declares in `IToolImplementation.OutboundData` where its arguments go: a chat which read from a mailbox keeps the tools whose data goes further than the mailbox allows from being offered and from running, see `ToolSelectionRules.IsOutboundDataAllowed`. A tool which brings content of a mailbox into the chat raises `ToolExecutionResult.RequiredOutboundDataRestriction`, next to `RequiredProviderConfidence` and `RequiredDataSecurity`. "Searching Mailboxes" in `documentation/Tools.md` explains the mail tools.
A tool which offers itself from the context of a chat instead of being selected, such as `semantic_search`, sets `Activation = ToolActivation.CONTEXT` and tailors its function to each request in `ResolveFunctionAsync`.
@@ -0,0 +1 @@
@AGENTS.md
+54
View File
@@ -0,0 +1,54 @@
# AGENTS.md
These instructions cover the Rust runtime in `runtime/`. They add to the `AGENTS.md` in the repository
root, which applies here as well.
## Rust Runtime (`runtime/`)
**Entry point:** `runtime/src/main.rs`
Key modules:
- `app_window.rs` - Tauri window management, updater integration
- `dotnet.rs` - Launches and manages the .NET sidecar process
- `runtime_api.rs` - Axum-based HTTPS API for .NET ↔ Rust communication
- `certificate.rs` - Generates self-signed TLS certificates for secure IPC
- `secret.rs` - Secure secret storage using OS keyring (Keychain/Credential Manager)
- `clipboard.rs` - Cross-platform clipboard operations
- `file_data.rs` - File processing for RAG (extracts text from PDF, DOCX, XLSX, PPTX, etc.)
- `encryption.rs` - AES-256-CBC encryption for sensitive data
- `pandoc.rs` - Integration with Pandoc for document conversion
- `log.rs` - Logging infrastructure using `flexi_logger`
**Every runtime API route requires the API token.** `require_api_token` in `runtime_api.rs` checks it
for all routes at once, so handlers take no `APIToken` argument. Register a new route in
`create_router`, before that call: `route_layer` protects only the routes registered before it, and
a route added afterward would be open to every process on the machine. The test
`every_route_requires_the_api_token` reads the routes from `create_router` and fails for any route
which answers without the token.
**Runtime API handlers never block.** All calls of the .NET app share one HTTP/2 connection, and the
task driving it may wait on exactly the Tokio worker which a blocking handler occupies, so a single
blocking handler can hold up the whole app. Therefore:
- File and disk access, the OS keyring, waiting for a `std::sync::Mutex`, and CPU-bound work such as
loading a tokenizer or scanning text run in `tokio::task::spawn_blocking`. Follow
`run_qdrant_edge_request` in `qdrant_edge_database.rs`, `prepare_image` and `prepare_image_sync` in
`image.rs`, or `sanitize_batch` in `prompt_injection/api.rs`.
- When an API offers a callback, as the file dialogs do, await the callback instead of calling the
blocking variant; see `await_dialog` in `file_actions.rs`.
- Short, bounded work, such as a single metadata lookup or writing a log line, may stay on the worker.
- When the blocking task fails, answer with an error, never with an empty value the app could take
for a valid answer.
- Never keep a `std::sync::MutexGuard` alive across an `.await`, not even with an explicit `drop`
before it: the future of the handler is then no longer `Send`, and Axum refuses it. Scope the guard
in a block instead.
- `runtime_api::test_support::assert_runtime_stays_free` tests a handler which waits for a lock. Run
such a test with `#[tokio::test(flavor = "multi_thread", worker_threads = 1)]`.
## IPC Communication Flow
1. Rust runtime starts and generates TLS certificate
2. Rust starts internal HTTPS API on random port
3. Rust launches .NET sidecar, passing: API port, certificate fingerprint, API token, secret key
4. .NET reads environment variables and establishes secure HTTPS connection to Rust
5. .NET requests an app port from Rust, starts Blazor Server on that port
6. Rust opens Tauri webview pointing to localhost:app_port
7. Bi-directional communication: .NET ↔ Rust via HTTPS API
+1
View File
@@ -0,0 +1 @@
@AGENTS.md