From 3a1f586f768f085bcdbbd07a685469d57accfa8e Mon Sep 17 00:00:00 2001 From: Thorsten Sommer Date: Sun, 11 Oct 2026 11:22:40 +0200 Subject: [PATCH] Split AGENTS.md into area guides so Codex reads it in full again (#1046) --- AGENTS.md | 386 ++---------------- app/AGENTS.md | 151 +++++++ app/Build/AGENTS.md | 51 +++ app/Build/CLAUDE.md | 1 + app/CLAUDE.md | 1 + app/MindWork AI Studio/Models/AGENTS.md | 15 + app/MindWork AI Studio/Models/CLAUDE.md | 1 + .../Tools/Databases/AGENTS.md | 33 ++ .../Tools/Databases/CLAUDE.md | 1 + .../Tools/PluginSystem/AGENTS.md | 26 ++ .../Tools/PluginSystem/CLAUDE.md | 1 + .../Tools/Services/Indexing/AGENTS.md | 47 +++ .../Tools/Services/Indexing/CLAUDE.md | 1 + .../Tools/ToolCallingSystem/AGENTS.md | 21 + .../Tools/ToolCallingSystem/CLAUDE.md | 1 + runtime/AGENTS.md | 54 +++ runtime/CLAUDE.md | 1 + 17 files changed, 435 insertions(+), 357 deletions(-) create mode 100644 app/AGENTS.md create mode 100644 app/Build/AGENTS.md create mode 100644 app/Build/CLAUDE.md create mode 100644 app/CLAUDE.md create mode 100644 app/MindWork AI Studio/Models/AGENTS.md create mode 100644 app/MindWork AI Studio/Models/CLAUDE.md create mode 100644 app/MindWork AI Studio/Tools/Databases/AGENTS.md create mode 100644 app/MindWork AI Studio/Tools/Databases/CLAUDE.md create mode 100644 app/MindWork AI Studio/Tools/PluginSystem/AGENTS.md create mode 100644 app/MindWork AI Studio/Tools/PluginSystem/CLAUDE.md create mode 100644 app/MindWork AI Studio/Tools/Services/Indexing/AGENTS.md create mode 100644 app/MindWork AI Studio/Tools/Services/Indexing/CLAUDE.md create mode 100644 app/MindWork AI Studio/Tools/ToolCallingSystem/AGENTS.md create mode 100644 app/MindWork AI Studio/Tools/ToolCallingSystem/CLAUDE.md create mode 100644 runtime/AGENTS.md create mode 100644 runtime/CLAUDE.md diff --git a/AGENTS.md b/AGENTS.md index df6565db..44b5ec93 100644 --- a/AGENTS.md +++ b/AGENTS.md @@ -1,6 +1,7 @@ # AGENTS.md -This file provides guidance to Claude Code (claude.ai/code) when working with code in this repository. +This file provides guidance to coding agents, such as Claude Code and Codex, when working with code in this +repository. ## Incremental implementation workflow @@ -13,6 +14,29 @@ a time. After each item: 4. Stop and wait until the developer has reviewed and committed the changes before continuing. 5. Never push the changes; the developer performs all pushes. +## Area guides + +Some areas of the code have instructions of their own, in an `AGENTS.md` next to the code. Claude Code +loads such a guide by itself once it works in that directory; other agents, Codex among them, do not. +Therefore, before you plan, read, or change anything in one of these areas, read its guide first: + +| Before you work on … | read first | +|---|---| +| any C#, Razor, or Lua code, or anything else under `app/` | `app/AGENTS.md` | +| anything under `runtime/` | `runtime/AGENTS.md` | +| a release, or crediting a contributor in a changelog entry or after a merge | `app/Build/AGENTS.md` | +| anything under `app/MindWork AI Studio/Tools/PluginSystem/`, or a new capability of configuration plugins | `app/MindWork AI Studio/Tools/PluginSystem/AGENTS.md` | +| anything under `app/MindWork AI Studio/Tools/ToolCallingSystem/` | `app/MindWork AI Studio/Tools/ToolCallingSystem/AGENTS.md` | +| anything under `app/MindWork AI Studio/Models/` or `app/Tests/Models/` | `app/MindWork AI Studio/Models/AGENTS.md` | +| anything under `app/MindWork AI Studio/Tools/Services/Indexing/`, or `DataSourceEmbeddingService` | `app/MindWork AI Studio/Tools/Services/Indexing/AGENTS.md` | +| anything under `app/MindWork AI Studio/Tools/Databases/` | `app/MindWork AI Studio/Tools/Databases/AGENTS.md` | + +This file keeps only what applies to the whole repository. `app/AGENTS.md` keeps what applies to all of the +.NET app, such as the rules for code which calls into one of the areas below it. Add instructions about a +single area to its guide instead. A new guide is an `AGENTS.md` plus a `CLAUDE.md` holding only +`@AGENTS.md`, listed in the table above. Never place one below `app/MindWork AI Studio/wwwroot/` or +`app/MindWork AI Studio/Plugins/`: both are embedded into the app and shipped to every user. + ## Working for an external contributor When you work for someone outside the core team, `CONTRIBUTING.md` applies in addition to this file. In @@ -105,344 +129,20 @@ through the IDE for the same reason they build there: mcp__rider__execute_terminal_command command: "cd app/Tests && dotnet test" ``` -An assembly-wide `[SetUpFixture]` in `app/Tests/TestHost.cs` fills the static application state that -the app itself only fills while starting up, `Program.LOGGER_FACTORY` above all. Types that -initialize a static logger from it — `Settings.Provider` among them — otherwise die in their type -initializer before the first assertion. Prefer writing new code so that it does not reach for such -statics at all. - The Rust tests run with `cargo test` in `runtime/`, through the `rustrover` MCP server. -## Architecture Details - -### Rust Runtime (`runtime/`) -**Entry point:** `runtime/src/main.rs` - -Key modules: -- `app_window.rs` - Tauri window management, updater integration -- `dotnet.rs` - Launches and manages the .NET sidecar process -- `runtime_api.rs` - Axum-based HTTPS API for .NET ↔ Rust communication -- `certificate.rs` - Generates self-signed TLS certificates for secure IPC -- `secret.rs` - Secure secret storage using OS keyring (Keychain/Credential Manager) -- `clipboard.rs` - Cross-platform clipboard operations -- `file_data.rs` - File processing for RAG (extracts text from PDF, DOCX, XLSX, PPTX, etc.) -- `encryption.rs` - AES-256-CBC encryption for sensitive data -- `pandoc.rs` - Integration with Pandoc for document conversion -- `log.rs` - Logging infrastructure using `flexi_logger` - -**Every runtime API route requires the API token.** `require_api_token` in `runtime_api.rs` checks it -for all routes at once, so handlers take no `APIToken` argument. Register a new route in -`create_router`, before that call: `route_layer` protects only the routes registered before it, and -a route added afterward would be open to every process on the machine. The test -`every_route_requires_the_api_token` reads the routes from `create_router` and fails for any route -which answers without the token. - -**Runtime API handlers never block.** All calls of the .NET app share one HTTP/2 connection, and the -task driving it may wait on exactly the Tokio worker which a blocking handler occupies, so a single -blocking handler can hold up the whole app. Therefore: - -- File and disk access, the OS keyring, waiting for a `std::sync::Mutex`, and CPU-bound work such as - loading a tokenizer or scanning text run in `tokio::task::spawn_blocking`. Follow - `run_qdrant_edge_request` in `qdrant_edge_database.rs`, `prepare_image` and `prepare_image_sync` in - `image.rs`, or `sanitize_batch` in `prompt_injection/api.rs`. -- When an API offers a callback, as the file dialogs do, await the callback instead of calling the - blocking variant; see `await_dialog` in `file_actions.rs`. -- Short, bounded work, such as a single metadata lookup or writing a log line, may stay on the worker. -- When the blocking task fails, answer with an error, never with an empty value the app could take - for a valid answer. -- Never keep a `std::sync::MutexGuard` alive across an `.await`, not even with an explicit `drop` - before it: the future of the handler is then no longer `Send`, and Axum refuses it. Scope the guard - in a block instead. -- `runtime_api::test_support::assert_runtime_stays_free` tests a handler which waits for a lock. Run - such a test with `#[tokio::test(flavor = "multi_thread", worker_threads = 1)]`. - -### .NET App (`app/MindWork AI Studio/`) -**Entry point:** `app/MindWork AI Studio/Program.cs` - -Key structure: -- **Program.cs** - Bootstraps Blazor Server, configures Kestrel, initializes encryption and Rust service -- **Provider/** - LLM provider implementations (OpenAI, Anthropic, Google, Mistral, etc.) - - `BaseProvider.cs` - Abstract base for all providers with streaming support - - `IProvider.cs` - Provider interface defining capabilities and streaming methods -- **Chat/** - Chat functionality and message handling -- **Assistants/** - Pre-configured assistants (translation, summarization, coding, etc.) - - `AssistantBase.razor` - Base component for all assistants -- **Agents/** - contains all agents, e.g., for data source selection, context validation, etc. - - `AgentDataSourceSelection.cs` - Selects appropriate data sources for queries - - `AgentRetrievalContextValidation.cs` - Validates retrieved context relevance -- **Tools/PluginSystem/** - Lua-based plugin system -- **Tools/Services/** - Core background services (settings, message bus, data sources, updates) -- **Tools/Rust/** - .NET wrapper for Rust API calls -- **Settings/** - Application settings and data models -- **Components/** - Reusable Blazor components -- **Pages/** - Top-level page components - -### IPC Communication Flow -1. Rust runtime starts and generates TLS certificate -2. Rust starts internal HTTPS API on random port -3. Rust launches .NET sidecar, passing: API port, certificate fingerprint, API token, secret key -4. .NET reads environment variables and establishes secure HTTPS connection to Rust -5. .NET requests an app port from Rust, starts Blazor Server on that port -6. Rust opens Tauri webview pointing to localhost:app_port -7. Bi-directional communication: .NET ↔ Rust via HTTPS API - -### Configuration and Metadata +## Configuration and Metadata - `metadata.txt` - Build metadata (version, build time, component versions) read by both Rust and .NET - `startup.env` - Development environment variables (generated by build script) - `.NET project` reads metadata.txt at build time and injects as assembly attributes -## Plugin System - -**Location:** `app/MindWork AI Studio/Plugins/` - -Plugins are written in Lua and provide: -- **Language plugins** - I18N translations (e.g., German language pack) -- **Configuration plugins** - Enterprise IT configurations for centrally managed providers, settings -- **Assistant plugins** - custom assistants and direct-chat launchers, subject to approval or a local security audit -- **Model plugins** - what an organization's own models can do, see `documentation/Models.md` - -**Example configuration plugin:** `app/MindWork AI Studio/Plugins/configuration/plugin.lua` - -Plugins can configure: -- Self-hosted LLM providers -- Update behavior -- Preview features visibility -- Preselected profiles -- Chat templates -- etc. - -Configuration plugins provide three kinds of values: -- **Managed settings:** simple values such as booleans, numbers, strings, enums, lists, or sets handled through `ManagedConfiguration`. These values may be locked or used as organization defaults. Which configuration plugin owns a locked setting is persisted in `Data.ManagedLockedConfigurations`, and organization defaults are tracked in `Data.ManagedEditableDefaults`. Both are cleaned up generically by `ManagedConfiguration.CleanupLeftOverManagedConfigurations(...)` when the owning plugin is gone. The value a setting had before a configuration plugin took it over is kept in `Data.ManagedUserValueSnapshots` and restored by that same clean-up, so removing a plugin hands the user's own value back instead of the app default. -- **Managed configuration objects:** complex Lua tables that are persisted into `SettingsManager.ConfigurationData`, implement `IConfigurationObject`, and are cleaned up through `PluginConfigurationObject.CleanLeftOverConfigurationObjects(...)`. Examples include providers, profiles, chat templates, data sources, and document analysis policies. -- **Live plugin content:** complex Lua tables that implement `ILivePluginContent` and are read live from running plugins instead of being persisted to `ConfigurationData`. Examples include `MANDATORY_INFOS` and `INTRODUCTIONS`. If live plugin content creates persistent side data, add a dedicated cleanup path for that side data, like mandatory-info acceptances. - -When adding configuration plugin capabilities: -- For managed settings, update the corresponding data class in `app/MindWork AI Studio/Settings/DataModel/` to call `ManagedConfiguration.Register(...)` and process the setting in `PluginConfiguration.TryProcessConfiguration`. Cleaning up the setting when its configuration plugin was removed needs no extra step: `ManagedConfiguration.CleanupLeftOverManagedConfigurations(...)` iterates all registered settings. Do not add per-setting cleanup calls to `PluginFactory.Loading.LoadAll`. -- For managed configuration objects, update `PluginConfigurationObject.cs` and `PluginConfigurationObjectType.cs`, persist them in the appropriate `ConfigurationData` collection, and add cleanup via `PluginConfigurationObject.CleanLeftOverConfigurationObjects(...)`. -- For live plugin content, add a data type implementing `ILivePluginContent`, parse it in `PluginConfiguration`, expose it through `PluginFactory`, and add any required cleanup only for persistent side data. -- Always document the new capability in `app/MindWork AI Studio/Plugins/configuration/plugin.lua`. - -## Tool Calling System - -**Documentation:** `documentation/Tools.md` - -When adding, changing, or removing model-driven tools, keep these parts in sync: -- `app/MindWork AI Studio/Tools/ToolCallingSystem/ToolCallingImplementations/` for the `IToolImplementation` class, which states its own `ToolDefinition` through `GetDefinition()`, written with `ToolSettingsSchemaBuilder` for its settings and `ToolParameterSchemaBuilder` for the arguments the model passes. There are no tool definition files; a tool arriving from elsewhere brings an `IToolDefinitionSource` instead. -- `app/MindWork AI Studio/Program.cs` for DI registration of the implementation. Registering it as an `IToolImplementation` is enough, because `CodeToolDefinitionSource` collects the definitions of all of them. -- An `IToolCollection` next to the tools, registered in `Program.cs`, when tools only make sense together. Selections, `DataTools.DisabledToolIds`, and the minimum provider confidence name tool collections; a tool outside a declared collection forms one under its own ID, and the ID of a tool inside one stands for its whole collection. Normalize a selection with `ToolRegistry.NormalizeSelection`, and expand it into tools with `ToolRegistry.ExpandSelection` only where the view of the model counts. -- `app/MindWork AI Studio/Tools/ToolCallingSystem/ToolSelectionRules.cs` when the shared tool-call limits change. A tool's own minimum provider confidence belongs in its definition, or in the definition of its collection, not here. -- `app/MindWork AI Studio/Tools/ToolCallingSystem/ToolSettingsOptionSources.cs` when a tool setting offers a fixed choice the app maintains, such as languages. Prefer this over spelling the values out in the settings schema; it keeps the list in one place and gives the user translated names. -- `app/MindWork AI Studio/Plugins/configuration/plugin.lua` to document each setting's field name, meaning, and data type. Tool settings need no code to be centrally manageable: an organization addresses them by `"."` in `DataTools.LockedToolSettings` or `DataTools.DefaultToolSettings`. - -Tool implementations must treat model-provided arguments as untrusted input. Validate settings and arguments, protect secrets with `SensitiveTraceArgumentNames`, use `ToolExecutionBlockedException` for intentional policy blocks, and check provider confidence before returning sensitive data to the model. - -A tool which belongs to a preview feature returns false from `IToolImplementation.IsAvailable` while the preview is switched off; the registry then leaves it out of every list, every request, and the token count, so no component has to check that preview for the tool. Every tool declares in `IToolImplementation.OutboundData` where its arguments go: a chat which read from a mailbox keeps the tools whose data goes further than the mailbox allows from being offered and from running, see `ToolSelectionRules.IsOutboundDataAllowed`. A tool which brings content of a mailbox into the chat raises `ToolExecutionResult.RequiredOutboundDataRestriction`, next to `RequiredProviderConfidence` and `RequiredDataSecurity`. "Searching Mailboxes" in `documentation/Tools.md` explains the mail tools. - -A tool which offers itself from the context of a chat instead of being selected, such as `semantic_search`, sets `Activation = ToolActivation.CONTEXT` and tailors its function to each request in `ResolveFunctionAsync`. Code which decides something on behalf of a request — whether the classic RAG process steps back, say — asks `ToolRegistry.GetOfferBlockReasonAsync` or `ToolRegistry.GetEffectiveRetrievalModeAsync` with the provider settings of the request (`IProvider.CreateSettingsProvider`), never a check of its own: two answers which drift apart leave a chat searching nothing or twice. - -## Model Capabilities - -**Documentation:** `documentation/Models.md` - -What a model can do is answered in `app/MindWork AI Studio/Models/`, through `provider.GetModelProfile()`. Never ask `ModelRegistry` directly from a component: the extension method is what adds the expert settings and what a provider's model list reported, and the registry alone answers neither. - -When adding, changing, or removing model knowledge, keep these parts in sync: -- `app/MindWork AI Studio/Models//.cs` for the family itself. Creating the class is enough — the source generator in `app/SourceGeneratedMappings/` collects every non-abstract `ModelFamily` and `IModelHost` at compile time, so there is no registration list. Do not add reflection here; `PublishTrimmed` is on. -- `app/Tests/Models/Corpus/` for the model IDs the family covers, marked as either unchanged or expected to change. A porting difference which nobody declared is what the corpus exists to catch. -- `app/MindWork AI Studio/Models/Kinds/` when the change is about what kind of model something is, rather than what it can do. These are ordinary rules of the same engine. -- `app/MindWork AI Studio/Models/Hosting/Hosts/` when a provider wraps model names or cannot pass an API through. A host unwraps and trims the transport; it states nothing about the model itself. -- `app/MindWork AI Studio/Plugins/models/plugin.lua` when a new field can be declared by an organization, and `app/MindWork AI Studio/Plugins/configuration/plugin.lua` when it can be overridden per provider instance. - -Rules are never tried in order: specificity is computed from the rule, and two rules of equal specificity on one name fail the test suite. State how a model reasons with `Reasoning(...)` — the three reasoning capabilities are override vocabulary and must never appear in a profile. Every family and every host has to name the page it was read from and the day somebody read it; `dotnet run verify-models` reports the ones which have gone stale. - -## RAG (Retrieval-Augmented Generation) - -RAG is available as a beta preview feature. Architecture: -- **External Retrieval Interface (ERI)** - Contract for integrating external data sources -- **Data Sources** - Local files and external data via ERI servers -- **Two ways to search** - By default, the chat model searches the data sources itself through the tool `semantic_search`, whenever a question calls for it. The classic process (`AISrcSelWithRetCtxVal`) searches them with every message instead, when the user chose so per chat (`DataSourceOptions.RetrievalMode`) or whenever the tool cannot be offered. `ToolRegistry.GetEffectiveRetrievalModeAsync` decides between the two; pass its answer to `DataSourceService`, because the agents only count as providers that see the data when they actually run. See "Searching Data Sources" in `documentation/Tools.md`. -- **Agents** - AI agents select data sources and validate retrieval quality, in the classic process only -- **Embedding providers** - Support for various embedding models -- **Vector database** - Qdrant Edge, embedded in the Rust runtime; see "Databases" below -- **Index database** - SQLite, holding the file fingerprints and the chunk texts for full-text search; see "Databases" below -- **File processing** - Extracts text from PDF, DOCX, XLSX via Rust runtime - -### Indexed data sources - -Everything AI Studio embeds itself runs through one pipeline in `app/MindWork AI Studio/Tools/Services/Indexing/`, -driven by `DataSourceEmbeddingService`, which queues the runs, prepares each one and owns the statuses. A new -kind of data source plugs into this pipeline instead of building its own. - -The parts: -- **`IIndexedDataSource`** (`Settings/`) - what indexing needs to know about a data source: confidence level, - embedding provider and chunk settings. `IDataSourceBase` is what every data source has. Implement - `IDataSource` on top only when classic RAG, Semantic Search and the agents should see the data source. A data - source kept in a list of its own implements `IIndexedDataSource` alone, and the compiler keeps it out of - `DataSources`. -- **`IIndexedSourceIndexer`** - one per kind of data source. `Supports` claims the data sources, `ProcessAsync` - finds and reads the documents of one run, and `TrackChanges` / `StopTracking` notice changes on their own. - `FileSourceIndexer` is the reference: a file system watcher per data source, fingerprints over path, size and - write time. -- **`IndexedRunContext`** - one prepared run: both stores, the embedding provider, the manifest and the - collection. `IndexDocumentAsync` embeds and stores one document; the cleanup methods remove what a failed - attempt left behind. -- **`EmbeddingDocument`** - one document: its key, its index row, its display name, and how to read its chunks. -- **`DocumentRunProgress`** - counts the documents, records indexed and failed ones in the stores, publishes the - status and completes the run. -- **`TextChunker`** - cuts text into chunks the embedding provider accepts. Pick one of its strategies; do not - write a chunker of your own. - -To add a kind of data source: -1. Write its indexer in `Tools/Services/Indexing/` and create it in `DataSourceEmbeddingService.CreateIndexers`, - which hands every indexer the same `TextChunker`. -2. Gate it in `IsSupportedIndexedSource`, behind a preview feature of its own while it is new. -3. When the data source is not kept in `DataSources`, add its list to `GetConfiguredIndexedSources`. Every lookup - by id and every pass over all data sources goes through it: the startup hash check, - `QueueAllInternalDataSourcesAsync` and `RefreshWatchers`. -4. Keep whatever the kind has to remember beyond its documents in tables of its own in the index store, added - by an EF Core migration (see "Databases"). -5. Report every status through `DocumentRunProgress`, so all rows of the embedding page behave alike. - -Rules which are easy to break: -- **A document key is not a path.** Only files use their full path as the key. Never pass a key through the - `Path` APIs: on Windows, `Path.GetFullPath` reads a key like `mail:…` as a file with an alternate data stream. -- **Ids and signature are pinned.** The formats in `IndexedDocumentIds` and the embedding signature - (`DataSourceEmbeddingService.BuildEmbeddingSignature`) are fixed by tests, because every stored chunk and every - index depends on them. When the metadata stored next to a chunk changes, raise `CHUNK_METADATA_VERSION` - deliberately: that rebuilds every index. -- **Every content path goes through the prompt injection filter.** Files pass the sanitizer of the runtime while - their text is extracted; a new kind of data source needs its own pass through `PromptInjectionGuardService`. -- **The service decides when, the indexer decides how.** Whether changes are tracked at all depends on the - automatic refresh setting and the startup hash check, and only the embedding service decides that. -- **Moving a `TB()` text into another class gives it a new I18N key**, so its translation is made anew during - the next localization run. - -## Databases - -Local RAG runs on two databases, addressed through `DatabaseRole`: - -- **`VECTOR_STORE`** — Qdrant Edge through the `qdrant-edge` crate, running **in-process inside the - Rust runtime**. There is no sidecar process, no port 6333 and no Qdrant API key; .NET reaches it - over the internal runtime API (`/system/qdrant-edge/*`, see `runtime/src/qdrant_edge_database.rs`), - secured by the same TLS and API token as every other runtime call. One store per data source, - named `rag_`, holding a single named vector `embedding` per point. -- **`INDEX_STORE`** — SQLite at `/databases/sqlite/rag-index.sqlite3`, reached - through EF Core. It holds the data sources, the file fingerprints, the chunk texts and an FTS5 - index over them, plus the files which permanently failed to index. - -`DatabaseClientProvider` is the only way to a client. It caches one per role and guards each role -with its own semaphore, so never construct a client yourself. - -When working on these, keep in mind: - -- **`GetDisplayInfo()` feeds the information page.** A new diagnostic value belongs in the client - that knows it, not in `Pages/Information.razor.cs`. The page renders whatever label-value pairs it - receives and stays free of per-database knowledge. -- **Let every probe in `GetDisplayInfo()` catch its own failure.** When the method throws, the page - replaces the *entire* block with the fallback client, so one unreadable value costs all the others - as well. -- **Raw SQL against SQLite goes through `context.Database.GetDbConnection()`**, not through - `SqlQueryRaw`: that one expects a column named `Value` and wraps the statement, so a `PRAGMA` - never works with it. -- **A new EF Core migration needs a `[DynamicDependency]`** in `IndexStoreSchemaMigrator`, because - `PublishTrimmed` is on and the migration type would otherwise be trimmed away. The "Schema version" - line on the information page shows the applied and pending counts, so a forgotten entry becomes - visible there. -- **Counts in the UI go through `long.CompactCount()` / `int.CompactCount()`** (`Tools/LongExtensions.cs`), - which shortens anything above 999 to `1.46k` or `4.51M` and formats it with the culture of the - active language plugin. Storage sizes are the exception: they keep using the byte formatters. - -## Enterprise IT Support - -AI Studio supports centralized configuration for enterprise environments: -- **Registry (Windows)** or **environment variables** (all platforms) specify configuration server URL and ID -- Configuration downloaded as ZIP containing Lua plugin -- Checks for updates every ~16 minutes via ETag -- Allows IT departments to pre-configure providers, settings, and chat templates - -**Documentation:** `documentation/Enterprise IT.md` - -## Provider Confidence System - -Multi-level confidence scheme allows users to control which providers see which data: -- Confidence levels: e.g. `NONE`, `LOW`, `MEDIUM`, `HIGH`, and some more granular levels -- Each assistant/feature can require a minimum confidence level -- Users assign confidence levels to providers based on trust - -**Implementation:** `app/MindWork AI Studio/Provider/Confidence.cs` - -## Dependencies and Frameworks - -**Rust:** -- Tauri 2 - Desktop application framework -- Axum - HTTPS API server -- tokio - Async runtime -- keyring - OS keyring integration -- pdfium-render - PDF text extraction -- calamine - Excel file parsing -- qdrant-edge - Embedded vector database - -**.NET:** -- Blazor Server - UI framework -- MudBlazor - Component library -- LuaCSharp - Lua scripting engine -- HtmlAgilityPack - HTML parsing -- ReverseMarkdown - HTML to Markdown conversion -- EF Core Sqlite + SQLitePCLRaw - the local RAG index - -## Security - -- **Encryption:** AES-256-CBC with PBKDF2 key derivation for sensitive data -- **IPC:** TLS-secured communication with random ports and API tokens -- **Secrets:** OS keyring for persistent secret storage (API keys, etc.) -- **Sandboxing:** Tauri provides OS-level sandboxing - -## Release Process - -1. Create changelog file: `app/MindWork AI Studio/wwwroot/changelog/vX.Y.Z.md` -2. Check that every contribution in the release is credited, see "Crediting contributors" below. The - authors of all pull requests merged since the last release are listed by - `gh pr list -R MindWorkAI/AI-Studio --state merged --limit 1000 --search "merged:>=" --json author --jq '[.[].author.login] | unique | .[]'`, - with the date of the last release tag from `git log -1 --format=%cs `; skip bots -3. Commit changelog -4. Run from `app/Build`: `dotnet run release --action ` -5. Create PR with version bump and changes -6. After PR merge, maintainer creates git tag: `vX.Y.Z` -7. GitHub Actions builds release binaries for all platforms -8. Binaries uploaded to GitHub Releases - -## Localization - -The app's texts are localized in two steps, and the developer always does the first one. - -1. The developer starts the app, which runs the I18N collector, and runs the localization assistant - in the app for German and US English. Agents never write these initial translations themselves: - they neither add nor regenerate entries in `app/MindWork AI Studio/Assistants/I18N/allTexts.lua`, - `app/MindWork AI Studio/Plugins/languages/en-us-97dfb1ba-50c4-4440-8dfa-6575daf543c8/plugin.lua`, - or `app/MindWork AI Studio/Plugins/languages/de-de-43065dbc-78d0-45b7-92be-f14c2926e2dc/plugin.lua`. - When new or changed texts are waiting for translation, remind the developer to start the app and - run the localization. -2. Afterward, agents always review the German translation. Compare the new and changed values of the - de-de `plugin.lua` with `main`, check them against the wording already established there, and - correct or improve them directly in that file. `allTexts.lua` and the en-us `plugin.lua` stay as - the assistant wrote them. - ## Important Development Notes - **File changes require Write/Edit tools** - Never use bash commands like `cat <` - **End of file formatting** - Do not append an extra empty line at the end of files. - **No automated formatting for Rust or .NET files** - Never run automated formatters on Rust files (`.rs`) or .NET files (`.cs`, `.razor`, `.csproj`, etc.). Only make the minimal manual formatting changes required for the specific edit. -- **I18N resources are generated** - The developer produces the translations by running the localization assistant in the app; agents only review and correct the German values afterward. See "Localization" above. - **Spaces in paths** - Always quote paths with spaces in bash commands -- **Agent-run builds** - Never start `.NET` or Rust builds in the agent's own shell; it is sandboxed. Use the `rider` and `rustrover` MCP servers instead, which build in the IDE outside that sandbox. See "Running builds from an agent" above. -- **Debug environment** - Reads `startup.env` file with IPC credentials -- **Production environment** - Runtime launches .NET sidecar with environment variables -- **MudBlazor** - Component library requires DI setup in Program.cs -- **Encryption** - Initialized before Rust service is marked ready -- **Message Bus** - Singleton event bus for cross-component communication inside the .NET app -- **Naming conventions** - Constants, enum members, and `static readonly` fields use `UPPER_SNAKE_CASE` such as `MY_CONSTANT`. - **Compatibility shims** - Temporary fallback or read-repair code must be documented in `documentation/compatibility-shims/` with an introduced date, remove-after date, code references, and removal checklist. Add a short code comment near the shim that references the document and remove-after date. Check this folder before adding similar fallback logic, and do not extend expired shims without explicit maintainer direction. Do not use this process for permanent settings schema migrations; those belong in `app/MindWork AI Studio/Settings/SettingsMigrations.cs`. -- **Empty lines** - Avoid adding extra empty lines at the end of files. ## Changelogs Changelogs are located in `app/MindWork AI Studio/wwwroot/changelog/` with filenames `vX.Y.Z.md`. These changelogs are meant to be for normal end-users @@ -473,34 +173,6 @@ inside the entry instead, even when that repeats a few words from another one. **Split a topic into several short entries** rather than growing a single long one, and address the reader with "you". -### Crediting contributors - -When a pull request is merged, or a release is prepared, check both places where we thank contributors. -This holds for everybody except the maintainer, members of the core team included, not only for external -contributors: - -- **The changelog entry of the change.** Thank the contributor at the end of the entry, in the form - `` (``)``, and call a first contribution out as such. -- **The "Code Contributions" list on the supporters page** in `app/MindWork AI Studio/Pages/Supporters.razor`. - Add contributors who are not listed yet, one - `` - each, at the end of the list. The list is sorted by the first merged contribution of each person; someone - whom the changelog thanks for work inside another person's pull request counts from that pull request on. - -The acknowledgment calls the person by their first name, unless the credit choice below rules out their -name, and says what they built. It never compares contributors or counts their contributions, and the -order of the list is chronological only. Bug reports -alone and commissioned work outside this repository get no entry on the supporters page; the changelog may -still thank them. - -The credit choice in the pull request template is binding: - -- **"Credit me with my GitHub username only":** use the GitHub username alone. No real name, neither in - the changelog nor in the acknowledgment text. -- **"Do not credit me":** neither a changelog mention nor an entry on the supporters page. -- **No choice ticked, or a pull request from before the template:** the GitHub username plus the name, when - the contributor shows it publicly on their GitHub profile (`gh api users/ --jq .name`) - or an earlier changelog already thanked them by that name. - -Acknowledgments on the supporters page are `T()` texts, so the two steps of "Localization" above apply: -remind the developer to run the localization, then review the German value. +**Thanking contributors.** Before an entry thanks a contributor, and whenever a pull request is merged, +read "Crediting contributors" in `app/Build/AGENTS.md` first: the credit choice in the pull request +template decides whether and how we name somebody. diff --git a/app/AGENTS.md b/app/AGENTS.md new file mode 100644 index 00000000..cc319ca0 --- /dev/null +++ b/app/AGENTS.md @@ -0,0 +1,151 @@ +# AGENTS.md + +These instructions cover the .NET app: all C#, Razor, and Lua code under `app/`, and everything else +there. They add to the `AGENTS.md` in the repository root, which applies here as well. Several areas below +`app/` have guides of their own; the table "Area guides" in the root `AGENTS.md` lists them. + +## .NET App (`app/MindWork AI Studio/`) +**Entry point:** `app/MindWork AI Studio/Program.cs` + +Key structure: +- **Program.cs** - Bootstraps Blazor Server, configures Kestrel, initializes encryption and Rust service +- **Provider/** - LLM provider implementations (OpenAI, Anthropic, Google, Mistral, etc.) + - `BaseProvider.cs` - Abstract base for all providers with streaming support + - `IProvider.cs` - Provider interface defining capabilities and streaming methods +- **Chat/** - Chat functionality and message handling +- **Assistants/** - Pre-configured assistants (translation, summarization, coding, etc.) + - `AssistantBase.razor` - Base component for all assistants +- **Agents/** - contains all agents, e.g., for data source selection, context validation, etc. + - `AgentDataSourceSelection.cs` - Selects appropriate data sources for queries + - `AgentRetrievalContextValidation.cs` - Validates retrieved context relevance +- **Tools/PluginSystem/** - Lua-based plugin system +- **Tools/Services/** - Core background services (settings, message bus, data sources, updates) +- **Tools/Rust/** - .NET wrapper for Rust API calls +- **Settings/** - Application settings and data models +- **Components/** - Reusable Blazor components +- **Pages/** - Top-level page components + +## Important Development Notes + +- **Naming conventions** - Constants, enum members, and `static readonly` fields use `UPPER_SNAKE_CASE` such as `MY_CONSTANT`. +- **Message Bus** - Singleton event bus for cross-component communication inside the .NET app +- **Encryption** - Initialized before Rust service is marked ready +- **Debug environment** - Reads `startup.env` file with IPC credentials +- **Production environment** - Runtime launches .NET sidecar with environment variables + +## Tests + +An assembly-wide `[SetUpFixture]` in `app/Tests/TestHost.cs` fills the static application state that +the app itself only fills while starting up, `Program.LOGGER_FACTORY` above all. Types that +initialize a static logger from it — `Settings.Provider` among them — otherwise die in their type +initializer before the first assertion. Prefer writing new code so that it does not reach for such +statics at all. + +## Localization + +The app's texts are localized in two steps, and the developer always does the first one. + +1. The developer starts the app, which runs the I18N collector, and runs the localization assistant + in the app for German and US English. Agents never write these initial translations themselves: + they neither add nor regenerate entries in `app/MindWork AI Studio/Assistants/I18N/allTexts.lua`, + `app/MindWork AI Studio/Plugins/languages/en-us-97dfb1ba-50c4-4440-8dfa-6575daf543c8/plugin.lua`, + or `app/MindWork AI Studio/Plugins/languages/de-de-43065dbc-78d0-45b7-92be-f14c2926e2dc/plugin.lua`. + When new or changed texts are waiting for translation, remind the developer to start the app and + run the localization. +2. Afterward, agents always review the German translation. Compare the new and changed values of the + de-de `plugin.lua` with `main`, check them against the wording already established there, and + correct or improve them directly in that file. `allTexts.lua` and the en-us `plugin.lua` stay as + the assistant wrote them. + +**Moving a `TB()` text into another class gives it a new I18N key**, so its translation is made anew during +the next localization run. + +## Plugin System + +**Location:** `app/MindWork AI Studio/Plugins/` + +Plugins are written in Lua and provide: +- **Language plugins** - I18N translations (e.g., German language pack) +- **Configuration plugins** - Enterprise IT configurations for centrally managed providers, settings +- **Assistant plugins** - custom assistants and direct-chat launchers, subject to approval or a local security audit +- **Model plugins** - what an organization's own models can do, see `documentation/Models.md` + +**Example configuration plugin:** `app/MindWork AI Studio/Plugins/configuration/plugin.lua` + +**Area guide:** `app/MindWork AI Studio/Tools/PluginSystem/AGENTS.md`, with the three kinds of values a +configuration plugin provides and how to add each of them. + +## Tool Calling System + +**Documentation:** `documentation/Tools.md` + +Selections, `DataTools.DisabledToolIds`, and the minimum provider confidence name tool collections; a tool outside a declared collection forms one under its own ID, and the ID of a tool inside one stands for its whole collection. Normalize a selection with `ToolRegistry.NormalizeSelection`, and expand it into tools with `ToolRegistry.ExpandSelection` only where the view of the model counts. + +Code which decides something on behalf of a request — whether the classic RAG process steps back, say — asks `ToolRegistry.GetOfferBlockReasonAsync` or `ToolRegistry.GetEffectiveRetrievalModeAsync` with the provider settings of the request (`IProvider.CreateSettingsProvider`), never a check of its own: two answers which drift apart leave a chat searching nothing or twice. + +**Area guide:** `app/MindWork AI Studio/Tools/ToolCallingSystem/AGENTS.md`, with how to add, change, or +remove a tool and the rules every tool implementation follows. + +## Model Capabilities + +**Documentation:** `documentation/Models.md` + +What a model can do is answered in `app/MindWork AI Studio/Models/`, through `provider.GetModelProfile()`. Never ask `ModelRegistry` directly from a component: the extension method is what adds the expert settings and what a provider's model list reported, and the registry alone answers neither. + +**Area guide:** `app/MindWork AI Studio/Models/AGENTS.md`, with how to add, change, or remove model +knowledge. + +## RAG (Retrieval-Augmented Generation) + +RAG is available as a beta preview feature. Architecture: +- **External Retrieval Interface (ERI)** - Contract for integrating external data sources +- **Data Sources** - Local files and external data via ERI servers +- **Two ways to search** - By default, the chat model searches the data sources itself through the tool `semantic_search`, whenever a question calls for it. The classic process (`AISrcSelWithRetCtxVal`) searches them with every message instead, when the user chose so per chat (`DataSourceOptions.RetrievalMode`) or whenever the tool cannot be offered. `ToolRegistry.GetEffectiveRetrievalModeAsync` decides between the two; pass its answer to `DataSourceService`, because the agents only count as providers that see the data when they actually run. See "Searching Data Sources" in `documentation/Tools.md`. +- **Agents** - AI agents select data sources and validate retrieval quality, in the classic process only +- **Embedding providers** - Support for various embedding models +- **Vector database** - Qdrant Edge, embedded in the Rust runtime; see "Databases" below +- **Index database** - SQLite, holding the file fingerprints and the chunk texts for full-text search; see "Databases" below +- **File processing** - Extracts text from PDF, DOCX, XLSX via Rust runtime + +### Indexed data sources + +Everything AI Studio embeds itself runs through one pipeline in `app/MindWork AI Studio/Tools/Services/Indexing/`, +driven by `DataSourceEmbeddingService`, which queues the runs, prepares each one and owns the statuses. A new +kind of data source plugs into this pipeline instead of building its own. + +**Every content path goes through the prompt injection filter.** Files pass the sanitizer of the runtime while +their text is extracted; a new kind of data source needs its own pass through `PromptInjectionGuardService`. + +**Area guide:** `app/MindWork AI Studio/Tools/Services/Indexing/AGENTS.md`, with the parts of the +pipeline, how to add a kind of data source, and the rules which are easy to break. + +## Databases + +`DatabaseClientProvider` is the only way to a client. It caches one per role and guards each role +with its own semaphore, so never construct a client yourself. + +**Counts in the UI go through `long.CompactCount()` / `int.CompactCount()`** (`Tools/LongExtensions.cs`), +which shortens anything above 999 to `1.46k` or `4.51M` and formats it with the culture of the +active language plugin. Storage sizes are the exception: they keep using the byte formatters. + +**Area guide:** `app/MindWork AI Studio/Tools/Databases/AGENTS.md`, with both stores and what to keep in +mind when working on them, EF Core migrations included. + +## Enterprise IT Support + +AI Studio supports centralized configuration for enterprise environments: +- **Registry (Windows)** or **environment variables** (all platforms) specify configuration server URL and ID +- Configuration downloaded as ZIP containing Lua plugin +- Checks for updates every ~16 minutes via ETag +- Allows IT departments to pre-configure providers, settings, and chat templates + +**Documentation:** `documentation/Enterprise IT.md` + +## Provider Confidence System + +Multi-level confidence scheme allows users to control which providers see which data: +- Confidence levels: e.g. `NONE`, `LOW`, `MEDIUM`, `HIGH`, and some more granular levels +- Each assistant/feature can require a minimum confidence level +- Users assign confidence levels to providers based on trust + +**Implementation:** `app/MindWork AI Studio/Provider/Confidence.cs` diff --git a/app/Build/AGENTS.md b/app/Build/AGENTS.md new file mode 100644 index 00000000..ec88fe5a --- /dev/null +++ b/app/Build/AGENTS.md @@ -0,0 +1,51 @@ +# AGENTS.md + +These instructions cover preparing a release and crediting contributors. They add to the `AGENTS.md` in +the repository root, which applies here as well; its "Changelogs" section explains how to write the +changelog entries. + +## Release Process + +1. Create changelog file: `app/MindWork AI Studio/wwwroot/changelog/vX.Y.Z.md` +2. Check that every contribution in the release is credited, see "Crediting contributors" below. The + authors of all pull requests merged since the last release are listed by + `gh pr list -R MindWorkAI/AI-Studio --state merged --limit 1000 --search "merged:>=" --json author --jq '[.[].author.login] | unique | .[]'`, + with the date of the last release tag from `git log -1 --format=%cs `; skip bots +3. Commit changelog +4. Run from `app/Build`: `dotnet run release --action ` +5. Create PR with version bump and changes +6. After PR merge, maintainer creates git tag: `vX.Y.Z` +7. GitHub Actions builds release binaries for all platforms +8. Binaries uploaded to GitHub Releases + +## Crediting contributors + +When a pull request is merged, or a release is prepared, check both places where we thank contributors. +This holds for everybody except the maintainer, members of the core team included, not only for external +contributors: + +- **The changelog entry of the change.** Thank the contributor at the end of the entry, in the form + `` (``)``, and call a first contribution out as such. +- **The "Code Contributions" list on the supporters page** in `app/MindWork AI Studio/Pages/Supporters.razor`. + Add contributors who are not listed yet, one + `` + each, at the end of the list. The list is sorted by the first merged contribution of each person; someone + whom the changelog thanks for work inside another person's pull request counts from that pull request on. + +The acknowledgment calls the person by their first name, unless the credit choice below rules out their +name, and says what they built. It never compares contributors or counts their contributions, and the +order of the list is chronological only. Bug reports +alone and commissioned work outside this repository get no entry on the supporters page; the changelog may +still thank them. + +The credit choice in the pull request template is binding: + +- **"Credit me with my GitHub username only":** use the GitHub username alone. No real name, neither in + the changelog nor in the acknowledgment text. +- **"Do not credit me":** neither a changelog mention nor an entry on the supporters page. +- **No choice ticked, or a pull request from before the template:** the GitHub username plus the name, when + the contributor shows it publicly on their GitHub profile (`gh api users/ --jq .name`) + or an earlier changelog already thanked them by that name. + +Acknowledgments on the supporters page are `T()` texts, so the two steps of "Localization" in the root +`AGENTS.md` apply: remind the developer to run the localization, then review the German value. diff --git a/app/Build/CLAUDE.md b/app/Build/CLAUDE.md new file mode 100644 index 00000000..eef4bd20 --- /dev/null +++ b/app/Build/CLAUDE.md @@ -0,0 +1 @@ +@AGENTS.md \ No newline at end of file diff --git a/app/CLAUDE.md b/app/CLAUDE.md new file mode 100644 index 00000000..eef4bd20 --- /dev/null +++ b/app/CLAUDE.md @@ -0,0 +1 @@ +@AGENTS.md \ No newline at end of file diff --git a/app/MindWork AI Studio/Models/AGENTS.md b/app/MindWork AI Studio/Models/AGENTS.md new file mode 100644 index 00000000..1ffab66b --- /dev/null +++ b/app/MindWork AI Studio/Models/AGENTS.md @@ -0,0 +1,15 @@ +# AGENTS.md + +These instructions cover the model knowledge in `app/MindWork AI Studio/Models/`. They add to the +`AGENTS.md` in the repository root, which applies here as well. + +## Model Capabilities + +When adding, changing, or removing model knowledge, keep these parts in sync: +- `app/MindWork AI Studio/Models//.cs` for the family itself. Creating the class is enough — the source generator in `app/SourceGeneratedMappings/` collects every non-abstract `ModelFamily` and `IModelHost` at compile time, so there is no registration list. Do not add reflection here; `PublishTrimmed` is on. +- `app/Tests/Models/Corpus/` for the model IDs the family covers, marked as either unchanged or expected to change. A porting difference which nobody declared is what the corpus exists to catch. +- `app/MindWork AI Studio/Models/Kinds/` when the change is about what kind of model something is, rather than what it can do. These are ordinary rules of the same engine. +- `app/MindWork AI Studio/Models/Hosting/Hosts/` when a provider wraps model names or cannot pass an API through. A host unwraps and trims the transport; it states nothing about the model itself. +- `app/MindWork AI Studio/Plugins/models/plugin.lua` when a new field can be declared by an organization, and `app/MindWork AI Studio/Plugins/configuration/plugin.lua` when it can be overridden per provider instance. + +Rules are never tried in order: specificity is computed from the rule, and two rules of equal specificity on one name fail the test suite. State how a model reasons with `Reasoning(...)` — the three reasoning capabilities are override vocabulary and must never appear in a profile. Every family and every host has to name the page it was read from and the day somebody read it; `dotnet run verify-models` reports the ones which have gone stale. diff --git a/app/MindWork AI Studio/Models/CLAUDE.md b/app/MindWork AI Studio/Models/CLAUDE.md new file mode 100644 index 00000000..eef4bd20 --- /dev/null +++ b/app/MindWork AI Studio/Models/CLAUDE.md @@ -0,0 +1 @@ +@AGENTS.md \ No newline at end of file diff --git a/app/MindWork AI Studio/Tools/Databases/AGENTS.md b/app/MindWork AI Studio/Tools/Databases/AGENTS.md new file mode 100644 index 00000000..fba40296 --- /dev/null +++ b/app/MindWork AI Studio/Tools/Databases/AGENTS.md @@ -0,0 +1,33 @@ +# AGENTS.md + +These instructions cover the databases in `app/MindWork AI Studio/Tools/Databases/`. They add to the +`AGENTS.md` in the repository root, which applies here as well. + +## Databases + +Local RAG runs on two databases, addressed through `DatabaseRole`: + +- **`VECTOR_STORE`** — Qdrant Edge through the `qdrant-edge` crate, running **in-process inside the + Rust runtime**. There is no sidecar process, no port 6333 and no Qdrant API key; .NET reaches it + over the internal runtime API (`/system/qdrant-edge/*`, see `runtime/src/qdrant_edge_database.rs`), + secured by the same TLS and API token as every other runtime call. One store per data source, + named `rag_`, holding a single named vector `embedding` per point. +- **`INDEX_STORE`** — SQLite at `/databases/sqlite/rag-index.sqlite3`, reached + through EF Core. It holds the data sources, the file fingerprints, the chunk texts and an FTS5 + index over them, plus the files which permanently failed to index. + +When working on these, keep in mind: + +- **`GetDisplayInfo()` feeds the information page.** A new diagnostic value belongs in the client + that knows it, not in `Pages/Information.razor.cs`. The page renders whatever label-value pairs it + receives and stays free of per-database knowledge. +- **Let every probe in `GetDisplayInfo()` catch its own failure.** When the method throws, the page + replaces the *entire* block with the fallback client, so one unreadable value costs all the others + as well. +- **Raw SQL against SQLite goes through `context.Database.GetDbConnection()`**, not through + `SqlQueryRaw`: that one expects a column named `Value` and wraps the statement, so a `PRAGMA` + never works with it. +- **A new EF Core migration needs a `[DynamicDependency]`** in `IndexStoreSchemaMigrator`, because + `PublishTrimmed` is on and the migration type would otherwise be trimmed away. The "Schema version" + line on the information page shows the applied and pending counts, so a forgotten entry becomes + visible there. diff --git a/app/MindWork AI Studio/Tools/Databases/CLAUDE.md b/app/MindWork AI Studio/Tools/Databases/CLAUDE.md new file mode 100644 index 00000000..eef4bd20 --- /dev/null +++ b/app/MindWork AI Studio/Tools/Databases/CLAUDE.md @@ -0,0 +1 @@ +@AGENTS.md \ No newline at end of file diff --git a/app/MindWork AI Studio/Tools/PluginSystem/AGENTS.md b/app/MindWork AI Studio/Tools/PluginSystem/AGENTS.md new file mode 100644 index 00000000..e311d358 --- /dev/null +++ b/app/MindWork AI Studio/Tools/PluginSystem/AGENTS.md @@ -0,0 +1,26 @@ +# AGENTS.md + +These instructions cover the plugin system in `app/MindWork AI Studio/Tools/PluginSystem/` and what +configuration plugins can provide. They add to the `AGENTS.md` in the repository root, which applies here +as well. + +## Configuration plugins + +Plugins can configure: +- Self-hosted LLM providers +- Update behavior +- Preview features visibility +- Preselected profiles +- Chat templates +- etc. + +Configuration plugins provide three kinds of values: +- **Managed settings:** simple values such as booleans, numbers, strings, enums, lists, or sets handled through `ManagedConfiguration`. These values may be locked or used as organization defaults. Which configuration plugin owns a locked setting is persisted in `Data.ManagedLockedConfigurations`, and organization defaults are tracked in `Data.ManagedEditableDefaults`. Both are cleaned up generically by `ManagedConfiguration.CleanupLeftOverManagedConfigurations(...)` when the owning plugin is gone. The value a setting had before a configuration plugin took it over is kept in `Data.ManagedUserValueSnapshots` and restored by that same clean-up, so removing a plugin hands the user's own value back instead of the app default. +- **Managed configuration objects:** complex Lua tables that are persisted into `SettingsManager.ConfigurationData`, implement `IConfigurationObject`, and are cleaned up through `PluginConfigurationObject.CleanLeftOverConfigurationObjects(...)`. Examples include providers, profiles, chat templates, data sources, and document analysis policies. +- **Live plugin content:** complex Lua tables that implement `ILivePluginContent` and are read live from running plugins instead of being persisted to `ConfigurationData`. Examples include `MANDATORY_INFOS` and `INTRODUCTIONS`. If live plugin content creates persistent side data, add a dedicated cleanup path for that side data, like mandatory-info acceptances. + +When adding configuration plugin capabilities: +- For managed settings, update the corresponding data class in `app/MindWork AI Studio/Settings/DataModel/` to call `ManagedConfiguration.Register(...)` and process the setting in `PluginConfiguration.TryProcessConfiguration`. Cleaning up the setting when its configuration plugin was removed needs no extra step: `ManagedConfiguration.CleanupLeftOverManagedConfigurations(...)` iterates all registered settings. Do not add per-setting cleanup calls to `PluginFactory.Loading.LoadAll`. +- For managed configuration objects, update `PluginConfigurationObject.cs` and `PluginConfigurationObjectType.cs`, persist them in the appropriate `ConfigurationData` collection, and add cleanup via `PluginConfigurationObject.CleanLeftOverConfigurationObjects(...)`. +- For live plugin content, add a data type implementing `ILivePluginContent`, parse it in `PluginConfiguration`, expose it through `PluginFactory`, and add any required cleanup only for persistent side data. +- Always document the new capability in `app/MindWork AI Studio/Plugins/configuration/plugin.lua`. diff --git a/app/MindWork AI Studio/Tools/PluginSystem/CLAUDE.md b/app/MindWork AI Studio/Tools/PluginSystem/CLAUDE.md new file mode 100644 index 00000000..eef4bd20 --- /dev/null +++ b/app/MindWork AI Studio/Tools/PluginSystem/CLAUDE.md @@ -0,0 +1 @@ +@AGENTS.md \ No newline at end of file diff --git a/app/MindWork AI Studio/Tools/Services/Indexing/AGENTS.md b/app/MindWork AI Studio/Tools/Services/Indexing/AGENTS.md new file mode 100644 index 00000000..b7d00c5b --- /dev/null +++ b/app/MindWork AI Studio/Tools/Services/Indexing/AGENTS.md @@ -0,0 +1,47 @@ +# AGENTS.md + +These instructions cover the indexing pipeline in `app/MindWork AI Studio/Tools/Services/Indexing/` and +`DataSourceEmbeddingService`. They add to the `AGENTS.md` in the repository root, which applies here as +well. + +## Indexed data sources + +The parts: +- **`IIndexedDataSource`** (`Settings/`) - what indexing needs to know about a data source: confidence level, + embedding provider and chunk settings. `IDataSourceBase` is what every data source has. Implement + `IDataSource` on top only when classic RAG, Semantic Search and the agents should see the data source. A data + source kept in a list of its own implements `IIndexedDataSource` alone, and the compiler keeps it out of + `DataSources`. +- **`IIndexedSourceIndexer`** - one per kind of data source. `Supports` claims the data sources, `ProcessAsync` + finds and reads the documents of one run, and `TrackChanges` / `StopTracking` notice changes on their own. + `FileSourceIndexer` is the reference: a file system watcher per data source, fingerprints over path, size and + write time. +- **`IndexedRunContext`** - one prepared run: both stores, the embedding provider, the manifest and the + collection. `IndexDocumentAsync` embeds and stores one document; the cleanup methods remove what a failed + attempt left behind. +- **`EmbeddingDocument`** - one document: its key, its index row, its display name, and how to read its chunks. +- **`DocumentRunProgress`** - counts the documents, records indexed and failed ones in the stores, publishes the + status and completes the run. +- **`TextChunker`** - cuts text into chunks the embedding provider accepts. Pick one of its strategies; do not + write a chunker of your own. + +To add a kind of data source: +1. Write its indexer in `Tools/Services/Indexing/` and create it in `DataSourceEmbeddingService.CreateIndexers`, + which hands every indexer the same `TextChunker`. +2. Gate it in `IsSupportedIndexedSource`, behind a preview feature of its own while it is new. +3. When the data source is not kept in `DataSources`, add its list to `GetConfiguredIndexedSources`. Every lookup + by id and every pass over all data sources goes through it: the startup hash check, + `QueueAllInternalDataSourcesAsync` and `RefreshWatchers`. +4. Keep whatever the kind has to remember beyond its documents in tables of its own in the index store, added + by an EF Core migration (see `app/MindWork AI Studio/Tools/Databases/AGENTS.md`). +5. Report every status through `DocumentRunProgress`, so all rows of the embedding page behave alike. + +Rules which are easy to break: +- **A document key is not a path.** Only files use their full path as the key. Never pass a key through the + `Path` APIs: on Windows, `Path.GetFullPath` reads a key like `mail:…` as a file with an alternate data stream. +- **Ids and signature are pinned.** The formats in `IndexedDocumentIds` and the embedding signature + (`DataSourceEmbeddingService.BuildEmbeddingSignature`) are fixed by tests, because every stored chunk and every + index depends on them. When the metadata stored next to a chunk changes, raise `CHUNK_METADATA_VERSION` + deliberately: that rebuilds every index. +- **The service decides when, the indexer decides how.** Whether changes are tracked at all depends on the + automatic refresh setting and the startup hash check, and only the embedding service decides that. diff --git a/app/MindWork AI Studio/Tools/Services/Indexing/CLAUDE.md b/app/MindWork AI Studio/Tools/Services/Indexing/CLAUDE.md new file mode 100644 index 00000000..eef4bd20 --- /dev/null +++ b/app/MindWork AI Studio/Tools/Services/Indexing/CLAUDE.md @@ -0,0 +1 @@ +@AGENTS.md \ No newline at end of file diff --git a/app/MindWork AI Studio/Tools/ToolCallingSystem/AGENTS.md b/app/MindWork AI Studio/Tools/ToolCallingSystem/AGENTS.md new file mode 100644 index 00000000..ef84081a --- /dev/null +++ b/app/MindWork AI Studio/Tools/ToolCallingSystem/AGENTS.md @@ -0,0 +1,21 @@ +# AGENTS.md + +These instructions cover the tool calling system in `app/MindWork AI Studio/Tools/ToolCallingSystem/`. They +add to the `AGENTS.md` in the repository root, which applies here as well; its "Tool Calling System" +section explains how selections name tool collections. + +## Tool Calling System + +When adding, changing, or removing model-driven tools, keep these parts in sync: +- `app/MindWork AI Studio/Tools/ToolCallingSystem/ToolCallingImplementations/` for the `IToolImplementation` class, which states its own `ToolDefinition` through `GetDefinition()`, written with `ToolSettingsSchemaBuilder` for its settings and `ToolParameterSchemaBuilder` for the arguments the model passes. There are no tool definition files; a tool arriving from elsewhere brings an `IToolDefinitionSource` instead. +- `app/MindWork AI Studio/Program.cs` for DI registration of the implementation. Registering it as an `IToolImplementation` is enough, because `CodeToolDefinitionSource` collects the definitions of all of them. +- An `IToolCollection` next to the tools, registered in `Program.cs`, when tools only make sense together. +- `app/MindWork AI Studio/Tools/ToolCallingSystem/ToolSelectionRules.cs` when the shared tool-call limits change. A tool's own minimum provider confidence belongs in its definition, or in the definition of its collection, not here. +- `app/MindWork AI Studio/Tools/ToolCallingSystem/ToolSettingsOptionSources.cs` when a tool setting offers a fixed choice the app maintains, such as languages. Prefer this over spelling the values out in the settings schema; it keeps the list in one place and gives the user translated names. +- `app/MindWork AI Studio/Plugins/configuration/plugin.lua` to document each setting's field name, meaning, and data type. Tool settings need no code to be centrally manageable: an organization addresses them by `"."` in `DataTools.LockedToolSettings` or `DataTools.DefaultToolSettings`. + +Tool implementations must treat model-provided arguments as untrusted input. Validate settings and arguments, protect secrets with `SensitiveTraceArgumentNames`, use `ToolExecutionBlockedException` for intentional policy blocks, and check provider confidence before returning sensitive data to the model. + +A tool which belongs to a preview feature returns false from `IToolImplementation.IsAvailable` while the preview is switched off; the registry then leaves it out of every list, every request, and the token count, so no component has to check that preview for the tool. Every tool declares in `IToolImplementation.OutboundData` where its arguments go: a chat which read from a mailbox keeps the tools whose data goes further than the mailbox allows from being offered and from running, see `ToolSelectionRules.IsOutboundDataAllowed`. A tool which brings content of a mailbox into the chat raises `ToolExecutionResult.RequiredOutboundDataRestriction`, next to `RequiredProviderConfidence` and `RequiredDataSecurity`. "Searching Mailboxes" in `documentation/Tools.md` explains the mail tools. + +A tool which offers itself from the context of a chat instead of being selected, such as `semantic_search`, sets `Activation = ToolActivation.CONTEXT` and tailors its function to each request in `ResolveFunctionAsync`. diff --git a/app/MindWork AI Studio/Tools/ToolCallingSystem/CLAUDE.md b/app/MindWork AI Studio/Tools/ToolCallingSystem/CLAUDE.md new file mode 100644 index 00000000..eef4bd20 --- /dev/null +++ b/app/MindWork AI Studio/Tools/ToolCallingSystem/CLAUDE.md @@ -0,0 +1 @@ +@AGENTS.md \ No newline at end of file diff --git a/runtime/AGENTS.md b/runtime/AGENTS.md new file mode 100644 index 00000000..61a61c28 --- /dev/null +++ b/runtime/AGENTS.md @@ -0,0 +1,54 @@ +# AGENTS.md + +These instructions cover the Rust runtime in `runtime/`. They add to the `AGENTS.md` in the repository +root, which applies here as well. + +## Rust Runtime (`runtime/`) +**Entry point:** `runtime/src/main.rs` + +Key modules: +- `app_window.rs` - Tauri window management, updater integration +- `dotnet.rs` - Launches and manages the .NET sidecar process +- `runtime_api.rs` - Axum-based HTTPS API for .NET ↔ Rust communication +- `certificate.rs` - Generates self-signed TLS certificates for secure IPC +- `secret.rs` - Secure secret storage using OS keyring (Keychain/Credential Manager) +- `clipboard.rs` - Cross-platform clipboard operations +- `file_data.rs` - File processing for RAG (extracts text from PDF, DOCX, XLSX, PPTX, etc.) +- `encryption.rs` - AES-256-CBC encryption for sensitive data +- `pandoc.rs` - Integration with Pandoc for document conversion +- `log.rs` - Logging infrastructure using `flexi_logger` + +**Every runtime API route requires the API token.** `require_api_token` in `runtime_api.rs` checks it +for all routes at once, so handlers take no `APIToken` argument. Register a new route in +`create_router`, before that call: `route_layer` protects only the routes registered before it, and +a route added afterward would be open to every process on the machine. The test +`every_route_requires_the_api_token` reads the routes from `create_router` and fails for any route +which answers without the token. + +**Runtime API handlers never block.** All calls of the .NET app share one HTTP/2 connection, and the +task driving it may wait on exactly the Tokio worker which a blocking handler occupies, so a single +blocking handler can hold up the whole app. Therefore: + +- File and disk access, the OS keyring, waiting for a `std::sync::Mutex`, and CPU-bound work such as + loading a tokenizer or scanning text run in `tokio::task::spawn_blocking`. Follow + `run_qdrant_edge_request` in `qdrant_edge_database.rs`, `prepare_image` and `prepare_image_sync` in + `image.rs`, or `sanitize_batch` in `prompt_injection/api.rs`. +- When an API offers a callback, as the file dialogs do, await the callback instead of calling the + blocking variant; see `await_dialog` in `file_actions.rs`. +- Short, bounded work, such as a single metadata lookup or writing a log line, may stay on the worker. +- When the blocking task fails, answer with an error, never with an empty value the app could take + for a valid answer. +- Never keep a `std::sync::MutexGuard` alive across an `.await`, not even with an explicit `drop` + before it: the future of the handler is then no longer `Send`, and Axum refuses it. Scope the guard + in a block instead. +- `runtime_api::test_support::assert_runtime_stays_free` tests a handler which waits for a lock. Run + such a test with `#[tokio::test(flavor = "multi_thread", worker_threads = 1)]`. + +## IPC Communication Flow +1. Rust runtime starts and generates TLS certificate +2. Rust starts internal HTTPS API on random port +3. Rust launches .NET sidecar, passing: API port, certificate fingerprint, API token, secret key +4. .NET reads environment variables and establishes secure HTTPS connection to Rust +5. .NET requests an app port from Rust, starts Blazor Server on that port +6. Rust opens Tauri webview pointing to localhost:app_port +7. Bi-directional communication: .NET ↔ Rust via HTTPS API diff --git a/runtime/CLAUDE.md b/runtime/CLAUDE.md new file mode 100644 index 00000000..eef4bd20 --- /dev/null +++ b/runtime/CLAUDE.md @@ -0,0 +1 @@ +@AGENTS.md \ No newline at end of file