Count the tool conversation in the token count (#973)
Build and Release / Read metadata (push) Blocked by required conditions
Build and Release / Sync Flatpak repo (push) Blocked by required conditions
Build and Release / Collect Flatpak artifacts (push) Blocked by required conditions
Build and Release / Verify (push) Waiting to run
Build and Release / Build app (${{ matrix.dotnet_runtime }}) (-aarch64-pc-windows-msvc.exe, win-arm64, windows-latest, aarch64-pc-windows-msvc, nsis,updater, nsis) (push) Blocked by required conditions
Build and Release / Build app (${{ matrix.dotnet_runtime }}) (-aarch64-unknown-linux-gnu, linux-arm64, ubuntu-22.04-arm, aarch64-unknown-linux-gnu, appimage,updater, appimage) (push) Blocked by required conditions
Build and Release / Build app (${{ matrix.dotnet_runtime }}) (-x86_64-apple-darwin, osx-x64, macos-latest, x86_64-apple-darwin, dmg,app,updater, dmg) (push) Blocked by required conditions
Build and Release / Build app (${{ matrix.dotnet_runtime }}) (-x86_64-pc-windows-msvc.exe, win-x64, windows-latest, x86_64-pc-windows-msvc, nsis,updater, nsis) (push) Blocked by required conditions
Build and Release / Build app (${{ matrix.dotnet_runtime }}) (-x86_64-unknown-linux-gnu, linux-x64, ubuntu-22.04, x86_64-unknown-linux-gnu, appimage,updater, appimage) (push) Blocked by required conditions
Build and Release / Prepare & create release (push) Blocked by required conditions
Build and Release / Publish release (push) Blocked by required conditions
Build and Release / Determine run mode (push) Waiting to run
Build and Release / Build app (${{ matrix.dotnet_runtime }}) (-aarch64-apple-darwin, osx-arm64, macos-latest, aarch64-apple-darwin, dmg,app,updater, dmg) (push) Blocked by required conditions

This commit is contained in:
Thorsten Sommer authored and GitHub committed 2026-09-14 19:54:23 +02:00
1 parent f869122070
commit 6ce7d856a3
15 files changed
+479 -58

No files matched your search

@@ -28,6 +28,15 @@ public sealed class AIJobService(SettingsManager settingsManager, MessageBus mes
public DateTimeOffset LastCheckpoint { get; set; }
/// <summary>
/// When the chat was last told that something happened which was not a streamed chunk.
/// </summary>
/// <remarks>
/// Kept on the job rather than in the loop which streams, because the tool calling reports
/// from outside that loop: it runs inside the provider call the loop is waiting on.
/// </remarks>
public DateTimeOffset LastActivityNotification { get; set; }
public bool IsCompletionStarted { get; set; }
public readonly Lock SyncRoot = new();
@@ -79,6 +88,44 @@ public sealed class AIJobService(SettingsManager settingsManager, MessageBus mes
return this.jobs.TryGetValue(jobId, out var job) ? job.ChatGenerationRequest?.ChatThread : null;
}
/// <summary>
/// Says that the answer of a chat has moved without a chunk having arrived.
/// </summary>
/// <remarks>
/// A model which calls tools asks several times before it says anything, and while it does,
/// this service sits in the provider call and hands nothing to the screen. But the request is
/// growing the whole time -- every tool result travels with the next round -- and the chat is
/// what recounts the tokens when it renders. Without this, the only thing which would ever ask
/// again is the ten-second heartbeat of the token tracker.
///
/// Throttled like the streamed chunks, and by the same setting: a round which calls five tools
/// in a row must not turn into five renders of the whole chat when somebody asked us to go easy
/// on their battery.
///
/// A chat without a running job is not an error. The same tool calling loop runs for the
/// assistants, which have no job behind them and no token count to update.
/// </remarks>
/// <param name="chatId">The chat whose answer moved.</param>
public async Task NotifyChatActivityAsync(Guid chatId)
{
if (!this.activeChatJobsByChatId.TryGetValue(chatId, out var jobId))
return;
if (!this.jobs.TryGetValue(jobId, out var job))
return;
lock (job.SyncRoot)
{
var now = DateTimeOffset.Now;
if (settingsManager.ConfigurationData.App.IsSavingEnergy && now - job.LastActivityNotification < STREAMING_EVENT_MIN_TIME)
return;
job.LastActivityNotification = now;
}
await this.NotifyChangedAsync(job);
}
public async Task<AIJobSnapshot?> TryStartChatGenerationAsync(ChatGenerationRequest request)
{
if (this.activeChatJobsByChatId.TryGetValue(request.ChatThread.ChatId, out var existingJobId))
@@ -309,6 +356,7 @@ public sealed class AIJobService(SettingsManager settingsManager, MessageBus mes
aiText.InitialRemoteWait = false;
aiText.IsStreaming = false;
aiText.Text = aiText.Text.RemoveThinkTags().Trim();
aiText.EndToolRun();
RemoveEmptyAIResponse(state);
@@ -49,4 +49,19 @@ public interface IToolCallingProviderAdapter
/// the others carry the failure in the content, which is where it has to be legible anyway.
/// </param>
public void RecordToolResult(string callId, string content, bool isError = false);
/// <summary>
/// The texts which everything recorded so far adds to the request of every following round.
/// </summary>
/// <remarks>
/// Kept by the adapter rather than by the loop, because the adapter is the only place which
/// knows what actually travels. The loop hands over arguments and results and would count
/// those; what the Responses API additionally demands back -- its reasoning items -- never
/// passes through the loop at all, and a conversation whose largest part is invisible is the
/// very thing this is here to rule out.<br/><br/>
/// These texts exist for as long as the adapter does, which is one streaming call. Nothing of
/// this reaches the next request the user sends: the accumulated conversation goes away with
/// the adapter.
/// </remarks>
public IReadOnlyList<string> RecordedRequestTexts { get; }
}
@@ -114,6 +114,7 @@ public sealed class ToolCallingLoop(ILogger<ToolCallingLoop> logger) : IToolCall
// The model's turn has to be recorded before its results, or the provider sees
// results for a turn it does not know about:
adapter.RecordAssistantTurn();
await context.PublishPendingToolConversationAsync(adapter);
foreach (var call in round.Calls)
{
@@ -124,6 +125,7 @@ public sealed class ToolCallingLoop(ILogger<ToolCallingLoop> logger) : IToolCall
toolResultCharacterCount += invalidContent.Length;
await context.AddToolInvocationAsync(invalidTrace);
adapter.RecordToolResult(call.CallId, invalidContent, isError: true);
await context.PublishPendingToolConversationAsync(adapter);
continue;
}
@@ -135,6 +137,7 @@ public sealed class ToolCallingLoop(ILogger<ToolCallingLoop> logger) : IToolCall
if (callsUnavailableInstruction is not null)
{
adapter.RecordToolResult(call.CallId, callsUnavailableInstruction);
await context.PublishPendingToolConversationAsync(adapter);
continue;
}
@@ -156,6 +159,7 @@ public sealed class ToolCallingLoop(ILogger<ToolCallingLoop> logger) : IToolCall
// A blocked call counts as a failure towards the model as much as an errored
// one does: in both cases it did not get the data it asked for.
adapter.RecordToolResult(call.CallId, toolContent, trace.Status is not ToolInvocationTraceStatus.SUCCESS);
await context.PublishPendingToolConversationAsync(adapter);
}
}
finally
@@ -1,5 +1,6 @@
using AIStudio.Chat;
using AIStudio.Provider;
using AIStudio.Tools.AIJobs;
namespace AIStudio.Tools.ToolCallingSystem.Harness;
@@ -55,7 +56,25 @@ public sealed class ToolCallingLoopContext
return;
this.CurrentAssistantContent.ToolInvocations.Add(trace);
await this.CurrentAssistantContent.StreamingEvent();
await this.AnnounceAsync(this.CurrentAssistantContent);
}
/// <summary>
/// Hands the conversation the adapter has accumulated to the assistant message.
/// </summary>
/// <remarks>
/// Called after every recording, not once per round: a round which reads five web pages is the
/// one during which the request grows the most, and a number which only moves between rounds
/// would stand still through exactly that.
/// </remarks>
/// <param name="adapter">The adapter of this run, which knows what it has recorded.</param>
public async Task PublishPendingToolConversationAsync(IToolCallingProviderAdapter adapter)
{
if (this.CurrentAssistantContent is null)
return;
this.CurrentAssistantContent.PendingToolConversation = [..adapter.RecordedRequestTexts];
await this.AnnounceAsync(this.CurrentAssistantContent);
}
/// <summary>
@@ -72,7 +91,7 @@ public sealed class ToolCallingLoopContext
ToolNames = [.. toolNames],
};
await this.CurrentAssistantContent.StreamingEvent();
await this.AnnounceAsync(this.CurrentAssistantContent);
}
/// <summary>
@@ -88,6 +107,32 @@ public sealed class ToolCallingLoopContext
return;
this.CurrentAssistantContent.ToolRuntimeStatus = new();
await this.CurrentAssistantContent.StreamingEvent();
await this.AnnounceAsync(this.CurrentAssistantContent);
}
/// <summary>
/// Says that something about the running answer has changed.
/// </summary>
/// <remarks>
/// Two receivers, because the screen is built from two of them. The content's own event
/// renders the message block, which is what shows a running tool and the calls it has made.
/// The job service renders the chat around it, and that is what recounts the tokens -- which
/// nothing else would ask for during a tool run: the chat hears about progress one streamed
/// chunk at a time, and a tool run produces none until it is over.<br/><br/>
/// One method rather than two calls at each of the four places above, because the second of
/// them is the one which is easy to forget.
/// </remarks>
/// <param name="content">The assistant message which changed.</param>
private async Task AnnounceAsync(ContentText content)
{
await content.StreamingEvent();
//
// Asked for here rather than taken as a dependency: the same loop runs for the assistants,
// where there is no job to tell and nothing which counts tokens.
//
var jobService = Program.SERVICE_PROVIDER.GetService<AIJobService>();
if (jobService is not null)
await jobService.NotifyChatActivityAsync(this.ChatThread.ChatId);
}
}