AI Powered Blazor WebAssembly: Client Side LLM Integration

Build offline capable AI UIs with Blazor WebAssembly. Learn LLM streaming, client side prompt engineering, and production patterns for type safe AI web apps.

0

Blazor WebAssembly lets you run .NET code directly in the browser. Add LLM integration to that equation, and you get something powerful: AI augmented web applications that don’t require heavy backend orchestration, that work offline, and that benefit from C# type safety and compiled performance.

This pattern is gaining traction. Teams are shipping production Blazor apps with real time AI capabilities built into the client. The approach solves real problems: reduced latency, lower API costs, better user experience during network hiccups, and faster iteration on AI first features.

Let’s walk through how to build it.

Why Client Side AI in Blazor Makes Sense

Traditional web apps route AI requests through a backend server. The browser sends a prompt, your API layer calls OpenAI or Anthropic, streams tokens back through WebSockets or Server Sent Events, and renders them in the UI. It works, but it adds latency, infrastructure cost, and operational complexity.

Blazor WebAssembly flips this. Your compiled C# code runs in the browser and calls LLM APIs directly. You save a round trip. You reduce backend load. You can cache responses locally in IndexedDB, so users get instant AI results even if the API is temporarily down or they’re on a slow connection.

Type safety is the second win. In JavaScript, a streaming LLM response is just JSON you parse and hope is correct. In C#, you define your prompt structure, response schema, and state management as strongly typed classes. The compiler catches mistakes before runtime. Your IDE gives you autocomplete. Refactoring is safe.

For enterprise teams, especially those already invested in .NET, this reduces cognitive overhead. You work in one language across frontend and backend. You write C#, deploy to the browser, and integrate with your existing backend services when you need to.

Setting Up the Foundation

Start with a new Blazor WebAssembly project:

dotnet new blazorwasm -n AiAssistantApp
cd AiAssistantApp

Add the packages you’ll need for HTTP client work and JSON serialization:

dotnet add package System.Net.Http.Json
dotnet add package System.Text.Json

For streaming LLM responses, you’ll want to parse Server Sent Events (SSE) or handle chunked transfer encoding. The HttpClient in .NET handles both, but you’ll wrap the response stream to process tokens as they arrive.

Define your core types. Here’s a minimal structure for an OpenAI style chat interface:

public class ChatMessage
{
    public string Role { get; set; } // "user", "assistant", "system"
    public string Content { get; set; }
}

public class ChatRequest
{
    public List<ChatMessage> Messages { get; set; } = new();
    public string Model { get; set; } = "gpt-4";
    public int MaxTokens { get; set; } = 1000;
    public float Temperature { get; set; } = 0.7f;
}

public class StreamedToken
{
    public string Delta { get; set; }
    public bool IsComplete { get; set; }
    public string ErrorMessage { get; set; }
}

These types give you a contract. Your UI components know exactly what to expect, and your API wrapper can validate responses before they reach your components.

Building the LLM Service Wrapper

Create a service that handles API calls and streaming. This is where the real work happens:

public class LlmService
{
    private readonly HttpClient _httpClient;
    private readonly string _apiKey;
    private readonly string _apiBaseUrl = "https://api.openai.com/v1";

    public LlmService(HttpClient httpClient, IConfiguration config)
    {
        _httpClient = httpClient;
        _apiKey = config["OpenAI:ApiKey"];
    }

    public async IAsyncEnumerable<StreamedToken> StreamChatCompletion(
        ChatRequest request,
        [EnumeratorCancellation] CancellationToken cancellationToken = default)
    {
        var url = $"{_apiBaseUrl}/chat/completions";
        var payload = new
        {
            messages = request.Messages,
            model = request.Model,
            max_tokens = request.MaxTokens,
            temperature = request.Temperature,
            stream = true
        };

        var content = new StringContent(
            JsonSerializer.Serialize(payload),
            Encoding.UTF8,
            "application/json");

        var requestMessage = new HttpRequestMessage(HttpMethod.Post, url)
        {
            Content = content,
            Headers = { { "Authorization", $"Bearer {_apiKey}" } }
        };

        using var response = await _httpClient.SendAsync(
            requestMessage,
            HttpCompletionOption.ResponseHeadersRead,
            cancellationToken);

        if (!response.IsSuccessStatusCode)
        {
            var errorContent = await response.Content.ReadAsStringAsync(cancellationToken);
            yield return new StreamedToken
            {
                IsComplete = true,
                ErrorMessage = $"API error: {response.StatusCode} - {errorContent}"
            };
            yield break;
        }

        using var stream = await response.Content.ReadAsStreamAsync(cancellationToken);
        using var reader = new StreamReader(stream);

        string line;
        while ((line = await reader.ReadLineAsync(cancellationToken)) != null)
        {
            if (string.IsNullOrWhiteSpace(line) || line == "[DONE]")
                continue;

            if (line.StartsWith("data: "))
            {
                var jsonData = line.Substring(6);
                try
                {
                    var chunk = JsonSerializer.Deserialize<OpenAiChunkResponse>(jsonData);
                    if (chunk?.Choices?.FirstOrDefault()?.Delta?.Content is string content)
                    {
                        yield return new StreamedToken { Delta = content };
                    }
                }
                catch (JsonException ex)
                {
                    yield return new StreamedToken
                    {
                        IsComplete = true,
                        ErrorMessage = $"Parse error: {ex.Message}"
                    };
                }
            }
        }

        yield return new StreamedToken { IsComplete = true };
    }
}

public class OpenAiChunkResponse
{
    [JsonPropertyName("choices")]
    public List<Choice> Choices { get; set; }
}

public class Choice
{
    [JsonPropertyName("delta")]
    public Delta Delta { get; set; }
}

public class Delta
{
    [JsonPropertyName("content")]
    public string Content { get; set; }
}

Key points here: we use HttpCompletionOption.ResponseHeadersRead to start reading the response body immediately, without buffering. We iterate through the stream line by line, parsing SSE format chunks. We yield each token as it arrives, so the UI can render it in real time.

Register this service in Program.cs:

builder.Services.AddScoped(sp => new HttpClient { BaseAddress = new Uri(builder.HostEnvironment.BaseAddress) });
builder.Services.AddScoped<LlmService>();

Building the UI Component

Now create a Blazor component that consumes the streaming service:

@page "/ai-assistant"
@using System.Collections.Generic
@inject LlmService LlmService
@implements IAsyncDisposable

<div class="chat-container">
    <div class="messages">
        @foreach (var msg in _messages)
        {
            <div class="message @msg.Role">
                <strong>@(msg.Role == "user" ? "You" : "Assistant"):</strong>
                <p>@msg.Content</p>
            </div>
        }
        @if (_isStreaming)
        {
            <div class="message assistant streaming">
                <strong>Assistant:</strong>
                <p>@_streamingContent<span class="cursor">|</span></p>
            </div>
        }
    </div>

    <div class="input-area">
        <textarea @bind="_userInput" placeholder="Ask me anything..." disabled="@_isStreaming"></textarea>
        <button @onclick="SendMessage" disabled="@(_isStreaming || string.IsNullOrWhiteSpace(_userInput))">Send</button>
    </div>

    @if (!string.IsNullOrEmpty(_errorMessage))
    {
        <div class="error-banner">@_errorMessage</div>
    }
</div>

@code {
    private List<ChatMessage> _messages = new();
    private string _userInput = "";
    private string _streamingContent = "";
    private string _errorMessage = "";
    private bool _isStreaming = false;
    private CancellationTokenSource _cancellationTokenSource;

    private async Task SendMessage()
    {
        if (string.IsNullOrWhiteSpace(_userInput))
            return;

        _errorMessage = "";
        var userMessage = new ChatMessage { Role = "user", Content = _userInput };
        _messages.Add(userMessage);
        _userInput = "";
        _streamingContent = "";
        _isStreaming = true;

        _cancellationTokenSource = new CancellationTokenSource();

        try
        {
            var request = new ChatRequest
            {
                Messages = _messages,
                Model = "gpt-4",
                Temperature = 0.7f
            };

            var assistantContent = new StringBuilder();

            await foreach (var token in LlmService.StreamChatCompletion(request, _cancellationTokenSource.Token))
            {
                if (token.IsComplete)
                {
                    if (!string.IsNullOrEmpty(token.ErrorMessage))
                    {
                        _errorMessage = token.ErrorMessage;
                    }
                    break;
                }

                if (!string.IsNullOrEmpty(token.Delta))
                {
                    assistantContent.Append(token.Delta);
                    _streamingContent = assistantContent.ToString();
                    StateHasChanged();
                }
            }

            _messages.Add(new ChatMessage { Role = "assistant", Content = assistantContent.ToString() });
        }
        catch (OperationCanceledException)
        {
            _errorMessage = "Request was cancelled.";
        }
        catch (Exception ex)
        {
            _errorMessage = $"Error: {ex.Message}";
        }
        finally
        {
            _isStreaming = false;
            _streamingContent = "";
        }
    }

    async ValueTask IAsyncDisposable.DisposeAsync()
    {
        _cancellationTokenSource?.Dispose();
    }
}

This component binds to the streaming service and renders tokens as they arrive. The UI updates in real time, giving users immediate visual feedback. We handle cancellation gracefully, and we store the full message in state once the stream completes.

Adding Offline Caching with IndexedDB

To make your AI assistant work offline, cache responses locally. Use the Blazor IndexedDB library:

dotnet add package TG.Blazor.IndexedDB

Create a cache service:

public class OfflineCacheService
{
    private readonly IDBFactory _idbFactory;
    private IDBDatabase _db;
    private const string DbName = "AiAssistantCache";
    private const string StoreName = "responses";

    public OfflineCacheService(IDBFactory idbFactory)
    {
        _idbFactory = idbFactory;
    }

    public async Task InitializeAsync()
    {
        var dbRequest = await _idbFactory.Open(DbName, version: 1);
        dbRequest.OnUpgradeNeeded += async e =>
        {
            var db = (IDBDatabase)e.NewVersion;
            if (!db.ObjectStoreNames.Contains(StoreName))
            {
                db.CreateObjectStore(StoreName, new ObjectStoreParameters { KeyPath = "id", AutoIncrement = true });
            }
        };
        _db = await dbRequest;
    }

    public async Task<string> GetCachedResponseAsync(string promptHash)
    {
        var transaction = _db.Transaction(new[] { StoreName }, "readonly");
        var store = transaction.ObjectStore(StoreName);
        var result = await store.Get(promptHash);
        return result?.GetProperty<string>("response");
    }

    public async Task CacheResponseAsync(string promptHash, string response)
    {
        var transaction = _db.Transaction(new[] { StoreName }, "readwrite");
        var store = transaction.ObjectStore(StoreName);
        await store.Add(new { id = promptHash, response = response, timestamp = DateTime.UtcNow });
    }
}

Before calling the LLM API, compute a hash of the prompt and check the cache. If you find a match, return it instantly. If not, fetch from the API and cache the result.

private string ComputePromptHash(string prompt)
{
    using var sha = System.Security.Cryptography.SHA256.Create();
    var hash = sha.ComputeHash(Encoding.UTF8.GetBytes(prompt));
    return Convert.ToBase64String(hash);
}

private async Task SendMessage()
{
    var promptHash = ComputePromptHash(_userInput);
    var cached = await _cacheService.GetCachedResponseAsync(promptHash);

    if (!string.IsNullOrEmpty(cached))
    {
        _messages.Add(new ChatMessage { Role = "assistant", Content = cached });
        return;
    }

    // Proceed with LLM API call as normal
}

Production Considerations

Exposing your API key in browser code is a security risk. Use a backend proxy or a service like Vercel’s AI SDK that handles credentials server side. Your Blazor app calls your own API, which forwards requests to OpenAI with the key safely stored in an environment variable.

Monitor token usage and implement rate limiting on the client. Track how many tokens each user consumes in a session. If you’re paying per token, this prevents runaway costs.

Implement graceful degradation. If the user is offline and the prompt isn’t in the cache, show a helpful message instead of a generic error. If the API is slow, show a loading state. If the network drops mid stream, let the user retry.

Test with realistic network conditions. Use Chrome DevTools throttling to simulate 3G, 4G, and offline scenarios. Streaming should feel responsive even on slow connections.

Consider splitting large conversations into separate IndexedDB entries. If a user has a 10 turn conversation, don’t serialize the entire history as one blob. Store each exchange separately so you can load and display them incrementally.

Performance Tuning

Blazor WebAssembly apps start with a download and initialization cost. The runtime, your compiled code, and dependencies are downloaded before your app runs. For an AI assistant, this is usually acceptable, but you can optimize:

Enable ahead of time (AOT) compilation in your .csproj to reduce download size and improve runtime performance:

<PropertyGroup>
    <PublishAot>true</PublishAot>
</PropertyGroup>

Use lazy loading for heavy components. If you have multiple AI features, load them on demand.

Minimize the number of state changes during streaming. Each call to StateHasChanged() triggers a render. Batch updates where possible, or use a virtualized list for very long conversations to render only visible messages.

Putting It Together

You now have the pieces: a typed LLM service that streams responses, a Blazor component that renders them in real time, offline caching with IndexedDB, and production ready error handling. The flow is straightforward:

User types a message. Component adds it to the conversation history. Service checks the cache. If not cached, it streams from the API. Each token updates the UI in real time. Once complete, the full response is stored locally and added to the message history. The user can continue the conversation, and the entire context is available for the next request.

What makes this powerful is that it’s all C#, all type safe, and all running in the browser. You’re not context switching between languages, you’re not managing two separate API contracts, and you’re not adding backend latency to every AI interaction. Your users get responsive AI features, and your team gets the productivity benefits of .NET.

This pattern is production ready today. The approach is documented and tested by the .NET community. If you’re building the next generation of enterprise web apps, especially in regulated industries where type safety and auditability matter, this is worth exploring.

Can I use Blazor WebAssembly for AI if my API key needs to stay secret?

Not directly. Never embed API keys in browser code. Instead, build a lightweight backend proxy. Your Blazor app calls your own API with a user token, and your backend forwards the request to OpenAI or Anthropic with your API key. This keeps credentials safe and lets you add logging, rate limiting, and cost controls server side.

Does offline caching work for all LLM use cases?

It works well for deterministic prompts: FAQs, documentation queries, templated responses. For creative or context dependent requests, caching is less useful because the same prompt might need different responses. You can still cache, but be aware that users might get stale or irrelevant results. Always show the cache timestamp so users know the response is cached.

How much does it cost to stream LLM responses from the browser?

You pay per token, same as calling the API from a server. The advantage is reduced latency and fewer backend round trips, which can save on intermediate processing costs. Implement client side token counting to show users the estimated cost before they submit a long prompt.

Will Blazor WebAssembly performance be acceptable with large conversations?

Yes, with caveats. Storing and rendering a 100 message conversation is fine. Rendering 1,000 messages will lag. Use virtualization (render only visible messages) or split old conversations into separate views. Blazor’s rendering is fast, but the browser still has limits. Test with your expected conversation size.

What’s the startup time for a Blazor WebAssembly app with AI features?

Startup time depends on your app size and network speed. Enable compression (gzip), use AOT compilation, and lazy load non critical components. For users on 3G or slower, consider a loading screen or skeleton UI. Cached assets load much faster on repeat visits.

Leave a Reply

Your email address will not be published. Required fields are marked *