Split AGENTS.md into area guides so Codex reads it in full again (#1046)

This commit is contained in:
Thorsten Sommer authored and GitHub committed 2026-10-11 11:22:40 +02:00
1 parent 8badadb35f
commit 3a1f586f76
17 files changed
+435 -357

No files matched your search

+151
View File
@@ -0,0 +1,151 @@
# AGENTS.md
These instructions cover the .NET app: all C#, Razor, and Lua code under `app/`, and everything else
there. They add to the `AGENTS.md` in the repository root, which applies here as well. Several areas below
`app/` have guides of their own; the table "Area guides" in the root `AGENTS.md` lists them.
## .NET App (`app/MindWork AI Studio/`)
**Entry point:** `app/MindWork AI Studio/Program.cs`
Key structure:
- **Program.cs** - Bootstraps Blazor Server, configures Kestrel, initializes encryption and Rust service
- **Provider/** - LLM provider implementations (OpenAI, Anthropic, Google, Mistral, etc.)
- `BaseProvider.cs` - Abstract base for all providers with streaming support
- `IProvider.cs` - Provider interface defining capabilities and streaming methods
- **Chat/** - Chat functionality and message handling
- **Assistants/** - Pre-configured assistants (translation, summarization, coding, etc.)
- `AssistantBase.razor` - Base component for all assistants
- **Agents/** - contains all agents, e.g., for data source selection, context validation, etc.
- `AgentDataSourceSelection.cs` - Selects appropriate data sources for queries
- `AgentRetrievalContextValidation.cs` - Validates retrieved context relevance
- **Tools/PluginSystem/** - Lua-based plugin system
- **Tools/Services/** - Core background services (settings, message bus, data sources, updates)
- **Tools/Rust/** - .NET wrapper for Rust API calls
- **Settings/** - Application settings and data models
- **Components/** - Reusable Blazor components
- **Pages/** - Top-level page components
## Important Development Notes
- **Naming conventions** - Constants, enum members, and `static readonly` fields use `UPPER_SNAKE_CASE` such as `MY_CONSTANT`.
- **Message Bus** - Singleton event bus for cross-component communication inside the .NET app
- **Encryption** - Initialized before Rust service is marked ready
- **Debug environment** - Reads `startup.env` file with IPC credentials
- **Production environment** - Runtime launches .NET sidecar with environment variables
## Tests
An assembly-wide `[SetUpFixture]` in `app/Tests/TestHost.cs` fills the static application state that
the app itself only fills while starting up, `Program.LOGGER_FACTORY` above all. Types that
initialize a static logger from it — `Settings.Provider` among them — otherwise die in their type
initializer before the first assertion. Prefer writing new code so that it does not reach for such
statics at all.
## Localization
The app's texts are localized in two steps, and the developer always does the first one.
1. The developer starts the app, which runs the I18N collector, and runs the localization assistant
in the app for German and US English. Agents never write these initial translations themselves:
they neither add nor regenerate entries in `app/MindWork AI Studio/Assistants/I18N/allTexts.lua`,
`app/MindWork AI Studio/Plugins/languages/en-us-97dfb1ba-50c4-4440-8dfa-6575daf543c8/plugin.lua`,
or `app/MindWork AI Studio/Plugins/languages/de-de-43065dbc-78d0-45b7-92be-f14c2926e2dc/plugin.lua`.
When new or changed texts are waiting for translation, remind the developer to start the app and
run the localization.
2. Afterward, agents always review the German translation. Compare the new and changed values of the
de-de `plugin.lua` with `main`, check them against the wording already established there, and
correct or improve them directly in that file. `allTexts.lua` and the en-us `plugin.lua` stay as
the assistant wrote them.
**Moving a `TB()` text into another class gives it a new I18N key**, so its translation is made anew during
the next localization run.
## Plugin System
**Location:** `app/MindWork AI Studio/Plugins/`
Plugins are written in Lua and provide:
- **Language plugins** - I18N translations (e.g., German language pack)
- **Configuration plugins** - Enterprise IT configurations for centrally managed providers, settings
- **Assistant plugins** - custom assistants and direct-chat launchers, subject to approval or a local security audit
- **Model plugins** - what an organization's own models can do, see `documentation/Models.md`
**Example configuration plugin:** `app/MindWork AI Studio/Plugins/configuration/plugin.lua`
**Area guide:** `app/MindWork AI Studio/Tools/PluginSystem/AGENTS.md`, with the three kinds of values a
configuration plugin provides and how to add each of them.
## Tool Calling System
**Documentation:** `documentation/Tools.md`
Selections, `DataTools.DisabledToolIds`, and the minimum provider confidence name tool collections; a tool outside a declared collection forms one under its own ID, and the ID of a tool inside one stands for its whole collection. Normalize a selection with `ToolRegistry.NormalizeSelection`, and expand it into tools with `ToolRegistry.ExpandSelection` only where the view of the model counts.
Code which decides something on behalf of a request — whether the classic RAG process steps back, say — asks `ToolRegistry.GetOfferBlockReasonAsync` or `ToolRegistry.GetEffectiveRetrievalModeAsync` with the provider settings of the request (`IProvider.CreateSettingsProvider`), never a check of its own: two answers which drift apart leave a chat searching nothing or twice.
**Area guide:** `app/MindWork AI Studio/Tools/ToolCallingSystem/AGENTS.md`, with how to add, change, or
remove a tool and the rules every tool implementation follows.
## Model Capabilities
**Documentation:** `documentation/Models.md`
What a model can do is answered in `app/MindWork AI Studio/Models/`, through `provider.GetModelProfile()`. Never ask `ModelRegistry` directly from a component: the extension method is what adds the expert settings and what a provider's model list reported, and the registry alone answers neither.
**Area guide:** `app/MindWork AI Studio/Models/AGENTS.md`, with how to add, change, or remove model
knowledge.
## RAG (Retrieval-Augmented Generation)
RAG is available as a beta preview feature. Architecture:
- **External Retrieval Interface (ERI)** - Contract for integrating external data sources
- **Data Sources** - Local files and external data via ERI servers
- **Two ways to search** - By default, the chat model searches the data sources itself through the tool `semantic_search`, whenever a question calls for it. The classic process (`AISrcSelWithRetCtxVal`) searches them with every message instead, when the user chose so per chat (`DataSourceOptions.RetrievalMode`) or whenever the tool cannot be offered. `ToolRegistry.GetEffectiveRetrievalModeAsync` decides between the two; pass its answer to `DataSourceService`, because the agents only count as providers that see the data when they actually run. See "Searching Data Sources" in `documentation/Tools.md`.
- **Agents** - AI agents select data sources and validate retrieval quality, in the classic process only
- **Embedding providers** - Support for various embedding models
- **Vector database** - Qdrant Edge, embedded in the Rust runtime; see "Databases" below
- **Index database** - SQLite, holding the file fingerprints and the chunk texts for full-text search; see "Databases" below
- **File processing** - Extracts text from PDF, DOCX, XLSX via Rust runtime
### Indexed data sources
Everything AI Studio embeds itself runs through one pipeline in `app/MindWork AI Studio/Tools/Services/Indexing/`,
driven by `DataSourceEmbeddingService`, which queues the runs, prepares each one and owns the statuses. A new
kind of data source plugs into this pipeline instead of building its own.
**Every content path goes through the prompt injection filter.** Files pass the sanitizer of the runtime while
their text is extracted; a new kind of data source needs its own pass through `PromptInjectionGuardService`.
**Area guide:** `app/MindWork AI Studio/Tools/Services/Indexing/AGENTS.md`, with the parts of the
pipeline, how to add a kind of data source, and the rules which are easy to break.
## Databases
`DatabaseClientProvider` is the only way to a client. It caches one per role and guards each role
with its own semaphore, so never construct a client yourself.
**Counts in the UI go through `long.CompactCount()` / `int.CompactCount()`** (`Tools/LongExtensions.cs`),
which shortens anything above 999 to `1.46k` or `4.51M` and formats it with the culture of the
active language plugin. Storage sizes are the exception: they keep using the byte formatters.
**Area guide:** `app/MindWork AI Studio/Tools/Databases/AGENTS.md`, with both stores and what to keep in
mind when working on them, EF Core migrations included.
## Enterprise IT Support
AI Studio supports centralized configuration for enterprise environments:
- **Registry (Windows)** or **environment variables** (all platforms) specify configuration server URL and ID
- Configuration downloaded as ZIP containing Lua plugin
- Checks for updates every ~16 minutes via ETag
- Allows IT departments to pre-configure providers, settings, and chat templates
**Documentation:** `documentation/Enterprise IT.md`
## Provider Confidence System
Multi-level confidence scheme allows users to control which providers see which data:
- Confidence levels: e.g. `NONE`, `LOW`, `MEDIUM`, `HIGH`, and some more granular levels
- Each assistant/feature can require a minimum confidence level
- Users assign confidence levels to providers based on trust
**Implementation:** `app/MindWork AI Studio/Provider/Confidence.cs`
+51
View File
@@ -0,0 +1,51 @@
# AGENTS.md
These instructions cover preparing a release and crediting contributors. They add to the `AGENTS.md` in
the repository root, which applies here as well; its "Changelogs" section explains how to write the
changelog entries.
## Release Process
1. Create changelog file: `app/MindWork AI Studio/wwwroot/changelog/vX.Y.Z.md`
2. Check that every contribution in the release is credited, see "Crediting contributors" below. The
authors of all pull requests merged since the last release are listed by
`gh pr list -R MindWorkAI/AI-Studio --state merged --limit 1000 --search "merged:>=<YYYY-MM-DD>" --json author --jq '[.[].author.login] | unique | .[]'`,
with the date of the last release tag from `git log -1 --format=%cs <last release tag>`; skip bots
3. Commit changelog
4. Run from `app/Build`: `dotnet run release --action <build|month|year>`
5. Create PR with version bump and changes
6. After PR merge, maintainer creates git tag: `vX.Y.Z`
7. GitHub Actions builds release binaries for all platforms
8. Binaries uploaded to GitHub Releases
## Crediting contributors
When a pull request is merged, or a release is prepared, check both places where we thank contributors.
This holds for everybody except the maintainer, members of the core team included, not only for external
contributors:
- **The changelog entry of the change.** Thank the contributor at the end of the entry, in the form
``<first name> <last name> (`<GitHub username>`)``, and call a first contribution out as such.
- **The "Code Contributions" list on the supporters page** in `app/MindWork AI Studio/Pages/Supporters.razor`.
Add contributors who are not listed yet, one
`<Supporter Name="<GitHub username>" Type="SupporterType.INDIVIDUAL" URL="https://github.com/<GitHub username>" Acknowledgment="@T("…")"/>`
each, at the end of the list. The list is sorted by the first merged contribution of each person; someone
whom the changelog thanks for work inside another person's pull request counts from that pull request on.
The acknowledgment calls the person by their first name, unless the credit choice below rules out their
name, and says what they built. It never compares contributors or counts their contributions, and the
order of the list is chronological only. Bug reports
alone and commissioned work outside this repository get no entry on the supporters page; the changelog may
still thank them.
The credit choice in the pull request template is binding:
- **"Credit me with my GitHub username only":** use the GitHub username alone. No real name, neither in
the changelog nor in the acknowledgment text.
- **"Do not credit me":** neither a changelog mention nor an entry on the supporters page.
- **No choice ticked, or a pull request from before the template:** the GitHub username plus the name, when
the contributor shows it publicly on their GitHub profile (`gh api users/<GitHub username> --jq .name`)
or an earlier changelog already thanked them by that name.
Acknowledgments on the supporters page are `T()` texts, so the two steps of "Localization" in the root
`AGENTS.md` apply: remind the developer to run the localization, then review the German value.
+1
View File
@@ -0,0 +1 @@
@AGENTS.md
+1
View File
@@ -0,0 +1 @@
@AGENTS.md
+15
View File
@@ -0,0 +1,15 @@
# AGENTS.md
These instructions cover the model knowledge in `app/MindWork AI Studio/Models/`. They add to the
`AGENTS.md` in the repository root, which applies here as well.
## Model Capabilities
When adding, changing, or removing model knowledge, keep these parts in sync:
- `app/MindWork AI Studio/Models/<Vendor>/<Family>.cs` for the family itself. Creating the class is enough — the source generator in `app/SourceGeneratedMappings/` collects every non-abstract `ModelFamily` and `IModelHost` at compile time, so there is no registration list. Do not add reflection here; `PublishTrimmed` is on.
- `app/Tests/Models/Corpus/` for the model IDs the family covers, marked as either unchanged or expected to change. A porting difference which nobody declared is what the corpus exists to catch.
- `app/MindWork AI Studio/Models/Kinds/` when the change is about what kind of model something is, rather than what it can do. These are ordinary rules of the same engine.
- `app/MindWork AI Studio/Models/Hosting/Hosts/` when a provider wraps model names or cannot pass an API through. A host unwraps and trims the transport; it states nothing about the model itself.
- `app/MindWork AI Studio/Plugins/models/plugin.lua` when a new field can be declared by an organization, and `app/MindWork AI Studio/Plugins/configuration/plugin.lua` when it can be overridden per provider instance.
Rules are never tried in order: specificity is computed from the rule, and two rules of equal specificity on one name fail the test suite. State how a model reasons with `Reasoning(...)` — the three reasoning capabilities are override vocabulary and must never appear in a profile. Every family and every host has to name the page it was read from and the day somebody read it; `dotnet run verify-models` reports the ones which have gone stale.
+1
View File
@@ -0,0 +1 @@
@AGENTS.md
@@ -0,0 +1,33 @@
# AGENTS.md
These instructions cover the databases in `app/MindWork AI Studio/Tools/Databases/`. They add to the
`AGENTS.md` in the repository root, which applies here as well.
## Databases
Local RAG runs on two databases, addressed through `DatabaseRole`:
- **`VECTOR_STORE`** — Qdrant Edge through the `qdrant-edge` crate, running **in-process inside the
Rust runtime**. There is no sidecar process, no port 6333 and no Qdrant API key; .NET reaches it
over the internal runtime API (`/system/qdrant-edge/*`, see `runtime/src/qdrant_edge_database.rs`),
secured by the same TLS and API token as every other runtime call. One store per data source,
named `rag_<data source guid>`, holding a single named vector `embedding` per point.
- **`INDEX_STORE`** — SQLite at `<data directory>/databases/sqlite/rag-index.sqlite3`, reached
through EF Core. It holds the data sources, the file fingerprints, the chunk texts and an FTS5
index over them, plus the files which permanently failed to index.
When working on these, keep in mind:
- **`GetDisplayInfo()` feeds the information page.** A new diagnostic value belongs in the client
that knows it, not in `Pages/Information.razor.cs`. The page renders whatever label-value pairs it
receives and stays free of per-database knowledge.
- **Let every probe in `GetDisplayInfo()` catch its own failure.** When the method throws, the page
replaces the *entire* block with the fallback client, so one unreadable value costs all the others
as well.
- **Raw SQL against SQLite goes through `context.Database.GetDbConnection()`**, not through
`SqlQueryRaw<T>`: that one expects a column named `Value` and wraps the statement, so a `PRAGMA`
never works with it.
- **A new EF Core migration needs a `[DynamicDependency]`** in `IndexStoreSchemaMigrator`, because
`PublishTrimmed` is on and the migration type would otherwise be trimmed away. The "Schema version"
line on the information page shows the applied and pending counts, so a forgotten entry becomes
visible there.
@@ -0,0 +1 @@
@AGENTS.md
@@ -0,0 +1,26 @@
# AGENTS.md
These instructions cover the plugin system in `app/MindWork AI Studio/Tools/PluginSystem/` and what
configuration plugins can provide. They add to the `AGENTS.md` in the repository root, which applies here
as well.
## Configuration plugins
Plugins can configure:
- Self-hosted LLM providers
- Update behavior
- Preview features visibility
- Preselected profiles
- Chat templates
- etc.
Configuration plugins provide three kinds of values:
- **Managed settings:** simple values such as booleans, numbers, strings, enums, lists, or sets handled through `ManagedConfiguration`. These values may be locked or used as organization defaults. Which configuration plugin owns a locked setting is persisted in `Data.ManagedLockedConfigurations`, and organization defaults are tracked in `Data.ManagedEditableDefaults`. Both are cleaned up generically by `ManagedConfiguration.CleanupLeftOverManagedConfigurations(...)` when the owning plugin is gone. The value a setting had before a configuration plugin took it over is kept in `Data.ManagedUserValueSnapshots` and restored by that same clean-up, so removing a plugin hands the user's own value back instead of the app default.
- **Managed configuration objects:** complex Lua tables that are persisted into `SettingsManager.ConfigurationData`, implement `IConfigurationObject`, and are cleaned up through `PluginConfigurationObject.CleanLeftOverConfigurationObjects(...)`. Examples include providers, profiles, chat templates, data sources, and document analysis policies.
- **Live plugin content:** complex Lua tables that implement `ILivePluginContent` and are read live from running plugins instead of being persisted to `ConfigurationData`. Examples include `MANDATORY_INFOS` and `INTRODUCTIONS`. If live plugin content creates persistent side data, add a dedicated cleanup path for that side data, like mandatory-info acceptances.
When adding configuration plugin capabilities:
- For managed settings, update the corresponding data class in `app/MindWork AI Studio/Settings/DataModel/` to call `ManagedConfiguration.Register(...)` and process the setting in `PluginConfiguration.TryProcessConfiguration`. Cleaning up the setting when its configuration plugin was removed needs no extra step: `ManagedConfiguration.CleanupLeftOverManagedConfigurations(...)` iterates all registered settings. Do not add per-setting cleanup calls to `PluginFactory.Loading.LoadAll`.
- For managed configuration objects, update `PluginConfigurationObject.cs` and `PluginConfigurationObjectType.cs`, persist them in the appropriate `ConfigurationData` collection, and add cleanup via `PluginConfigurationObject.CleanLeftOverConfigurationObjects(...)`.
- For live plugin content, add a data type implementing `ILivePluginContent`, parse it in `PluginConfiguration`, expose it through `PluginFactory`, and add any required cleanup only for persistent side data.
- Always document the new capability in `app/MindWork AI Studio/Plugins/configuration/plugin.lua`.
@@ -0,0 +1 @@
@AGENTS.md
@@ -0,0 +1,47 @@
# AGENTS.md
These instructions cover the indexing pipeline in `app/MindWork AI Studio/Tools/Services/Indexing/` and
`DataSourceEmbeddingService`. They add to the `AGENTS.md` in the repository root, which applies here as
well.
## Indexed data sources
The parts:
- **`IIndexedDataSource`** (`Settings/`) - what indexing needs to know about a data source: confidence level,
embedding provider and chunk settings. `IDataSourceBase` is what every data source has. Implement
`IDataSource` on top only when classic RAG, Semantic Search and the agents should see the data source. A data
source kept in a list of its own implements `IIndexedDataSource` alone, and the compiler keeps it out of
`DataSources`.
- **`IIndexedSourceIndexer`** - one per kind of data source. `Supports` claims the data sources, `ProcessAsync`
finds and reads the documents of one run, and `TrackChanges` / `StopTracking` notice changes on their own.
`FileSourceIndexer` is the reference: a file system watcher per data source, fingerprints over path, size and
write time.
- **`IndexedRunContext`** - one prepared run: both stores, the embedding provider, the manifest and the
collection. `IndexDocumentAsync` embeds and stores one document; the cleanup methods remove what a failed
attempt left behind.
- **`EmbeddingDocument`** - one document: its key, its index row, its display name, and how to read its chunks.
- **`DocumentRunProgress`** - counts the documents, records indexed and failed ones in the stores, publishes the
status and completes the run.
- **`TextChunker`** - cuts text into chunks the embedding provider accepts. Pick one of its strategies; do not
write a chunker of your own.
To add a kind of data source:
1. Write its indexer in `Tools/Services/Indexing/` and create it in `DataSourceEmbeddingService.CreateIndexers`,
which hands every indexer the same `TextChunker`.
2. Gate it in `IsSupportedIndexedSource`, behind a preview feature of its own while it is new.
3. When the data source is not kept in `DataSources`, add its list to `GetConfiguredIndexedSources`. Every lookup
by id and every pass over all data sources goes through it: the startup hash check,
`QueueAllInternalDataSourcesAsync` and `RefreshWatchers`.
4. Keep whatever the kind has to remember beyond its documents in tables of its own in the index store, added
by an EF Core migration (see `app/MindWork AI Studio/Tools/Databases/AGENTS.md`).
5. Report every status through `DocumentRunProgress`, so all rows of the embedding page behave alike.
Rules which are easy to break:
- **A document key is not a path.** Only files use their full path as the key. Never pass a key through the
`Path` APIs: on Windows, `Path.GetFullPath` reads a key like `mail:…` as a file with an alternate data stream.
- **Ids and signature are pinned.** The formats in `IndexedDocumentIds` and the embedding signature
(`DataSourceEmbeddingService.BuildEmbeddingSignature`) are fixed by tests, because every stored chunk and every
index depends on them. When the metadata stored next to a chunk changes, raise `CHUNK_METADATA_VERSION`
deliberately: that rebuilds every index.
- **The service decides when, the indexer decides how.** Whether changes are tracked at all depends on the
automatic refresh setting and the startup hash check, and only the embedding service decides that.
@@ -0,0 +1 @@
@AGENTS.md
@@ -0,0 +1,21 @@
# AGENTS.md
These instructions cover the tool calling system in `app/MindWork AI Studio/Tools/ToolCallingSystem/`. They
add to the `AGENTS.md` in the repository root, which applies here as well; its "Tool Calling System"
section explains how selections name tool collections.
## Tool Calling System
When adding, changing, or removing model-driven tools, keep these parts in sync:
- `app/MindWork AI Studio/Tools/ToolCallingSystem/ToolCallingImplementations/` for the `IToolImplementation` class, which states its own `ToolDefinition` through `GetDefinition()`, written with `ToolSettingsSchemaBuilder` for its settings and `ToolParameterSchemaBuilder` for the arguments the model passes. There are no tool definition files; a tool arriving from elsewhere brings an `IToolDefinitionSource` instead.
- `app/MindWork AI Studio/Program.cs` for DI registration of the implementation. Registering it as an `IToolImplementation` is enough, because `CodeToolDefinitionSource` collects the definitions of all of them.
- An `IToolCollection` next to the tools, registered in `Program.cs`, when tools only make sense together.
- `app/MindWork AI Studio/Tools/ToolCallingSystem/ToolSelectionRules.cs` when the shared tool-call limits change. A tool's own minimum provider confidence belongs in its definition, or in the definition of its collection, not here.
- `app/MindWork AI Studio/Tools/ToolCallingSystem/ToolSettingsOptionSources.cs` when a tool setting offers a fixed choice the app maintains, such as languages. Prefer this over spelling the values out in the settings schema; it keeps the list in one place and gives the user translated names.
- `app/MindWork AI Studio/Plugins/configuration/plugin.lua` to document each setting's field name, meaning, and data type. Tool settings need no code to be centrally manageable: an organization addresses them by `"<toolId>.<fieldName>"` in `DataTools.LockedToolSettings` or `DataTools.DefaultToolSettings`.
Tool implementations must treat model-provided arguments as untrusted input. Validate settings and arguments, protect secrets with `SensitiveTraceArgumentNames`, use `ToolExecutionBlockedException` for intentional policy blocks, and check provider confidence before returning sensitive data to the model.
A tool which belongs to a preview feature returns false from `IToolImplementation.IsAvailable` while the preview is switched off; the registry then leaves it out of every list, every request, and the token count, so no component has to check that preview for the tool. Every tool declares in `IToolImplementation.OutboundData` where its arguments go: a chat which read from a mailbox keeps the tools whose data goes further than the mailbox allows from being offered and from running, see `ToolSelectionRules.IsOutboundDataAllowed`. A tool which brings content of a mailbox into the chat raises `ToolExecutionResult.RequiredOutboundDataRestriction`, next to `RequiredProviderConfidence` and `RequiredDataSecurity`. "Searching Mailboxes" in `documentation/Tools.md` explains the mail tools.
A tool which offers itself from the context of a chat instead of being selected, such as `semantic_search`, sets `Activation = ToolActivation.CONTEXT` and tailors its function to each request in `ResolveFunctionAsync`.
@@ -0,0 +1 @@
@AGENTS.md