View on GitHub

CS357

Foundations of Artificial Intelligence

CS357: Foundations of Artificial Intelligence - Building an AI Chess Coach

Purpose

To dissect a complete working web app and see exactly how a language model gets wired into real software through API calls.

About This Tutorial

This tutorial dissects a complete, working web app (the Chess AI Coach) to show exactly how a language model gets wired into real software through API calls. We move from what the app is → the three-layer architecture → one function that talks to three different providers → prompt engineering and structured JSON output for coaching → keeping your API keys safe → wiring the AI into the user interface.

The app is a single self-contained HTML file. You can open it, read every line, and change it. Everything you learned in the RESTful LLM Access activity (the /v1/chat/completions payload, the choices[0].message.content response path, provider portability) reappears here, this time in JavaScript running inside a browser instead of Python in a notebook.

Key Concepts

Term Plain-English Definition Where You’ll Meet It
Single-file app An entire application (markup, styles, logic) in one .html file that runs by opening it in a browser chess-ai-coach.html loads React and Babel from a CDN and needs no build step
fetch The browser’s built-in function for making HTTP requests from JavaScript; the front-end equivalent of Python’s requests await fetch("https://api.anthropic.com/v1/messages", { method: "POST", ... })
Provider-agnostic dispatcher One function that accepts a prompt and routes it to whichever AI provider the user selected, hiding the per-provider differences callTextModel(providerConfig, prompt) handles Anthropic, OpenAI, and Open WebUI
System vs. user content The instruction framing (“you are a chess coach”) versus the specific request (this position, this move) The coaching prompt states the coach’s role, then supplies the FEN and move
Structured output Asking the model to answer as parseable JSON instead of prose, so the program can use the value {"eval": 0.5} for the evaluation bar; {"elo": 1200, "label": "Intermediate"}
FEN / SAN Standard text encodings of a chess position (FEN) and a single move (SAN), the tokens we hand the model rnbqkbnr/pppppppp/... w KQkq - 0 1; Nf3, O-O, exd5
Client-side key exposure The risk that an API key placed in browser code is visible to anyone using or inspecting that browser The key shows up in the browser’s Network tab on every request
Backend proxy A small server you own that holds the secret key and forwards browser requests to the provider Browser -> your /api/coach endpoint -> provider; the key never leaves your server
Graceful degradation The app stays fully usable when the optional AI is unavailable With no provider configured, the board still plays against a local engine

Before You Start

What you need: Ollama running locally. Section 0 lets you try the finished app before building anything.

What you will have at the end: a working chess coach that explains moves, built from API calls you can read end to end.

Take the sections in order, since each builds on the one before it. Run the code blocks as you come to them instead of reading past them.


0. Try the App First

Before reading any code, play the app so the rest of the activity has something concrete to attach to.

  1. Download or open chess-ai-coach.html.
  2. Because browsers restrict fetch from file:// pages, serve the folder with a tiny local web server and open it over http://:

Code Cell


# From the folder that contains chess-ai-coach.html
python -m http.server 8000

# then open http://localhost:8000/chess-ai-coach.html
  1. Leave the provider set to Local only and play a few moves. The board enforces full chess rules and a computer opponent replies, with no AI and no API key at all. Keep that in mind: the language model is an addition, not the engine.

Part I: What We’re Building and the Three-Layer Architecture

In this part you build a mental map of the app before touching the AI. The single most useful habit when adding AI to software is knowing which parts are deterministic program logic and which parts call the model, because they fail, cost, and get tested in completely different ways.

1. The App at a Glance

The Chess AI Coach lets you play a full game against a built-in computer opponent, and (if you connect a language model) it adds a coaching layer on top:

Every one of those features is optional. The board, the legal-move enforcement, the computer opponent, and a fallback evaluation all run locally in the browser with no network calls.

2. Three Layers, Cleanly Separated

The file is organized into three layers, and the separation is deliberate:

\[\underbrace{\text{Chess Engine}}_{\text{pure logic}} \;\;\rightarrow\;\; \underbrace{\text{React UI}}_{\text{display + input}} \;\;\leftarrow\;\; \underbrace{\text{AI Layer}}_{\text{API calls}}\]
Layer What it does Representative functions Needs the network?
Chess engine Enforces the rules, generates legal moves, plays the computer’s move, scores a position initialState, legalMoves, applyMove, minimax, computerMove, evaluate, boardToFEN, moveToSAN, buildPGN No, pure, deterministic
React UI Draws the board, handles clicks and drags, shows commentary and meters the ChessAICoach component and its useState/useEffect hooks No
AI layer Turns a position into a prompt, calls a provider, parses the reply callTextModel, getAICommentary, getAIEvaluation, getSideAIElo Yes, this is where the API calls live

The engine is the kind of code you can unit-test exhaustively: given this board, legalMoves must return exactly these moves. The AI layer is the opposite: it calls a probabilistic model over the network, so it can be slow, cost money, fail, or return something unexpected. Keeping them apart means a bug in the coach can never make the board illegal, and the game never depends on a server being reachable.


Locating the Seams

Open chess-ai-coach.html and skim the top-level function names. Group them into the three layers above.

Questions to Work Through

  1. evaluate(state) (local, in the engine) and getAIEvaluation(providerConfig, fen) (in the AI layer) both produce a single number describing who is winning. Why does the app keep both, and when does each one run?

    Hint: Look at the useEffect that sets evalScore. When no provider is configured (!aiEnabled), it uses evaluate(gameState) / 100. When a provider is configured, it calls getAIEvaluation and falls back to evaluate if the call throws. One is free, instant, and deterministic; the other needs a working API call.

  2. The function boardToFEN(state) belongs to the engine, but the AI layer depends on it heavily. What is the FEN string for, and why is it the natural thing to send to a language model instead of, say, the raw 8×8 JavaScript array?

Hint: FEN is a compact, standard, text encoding of a position. Models have seen enormous amounts of FEN in training; a nested JS array is not something a model reads fluently, and it would waste tokens. Text-in, text-out; the engine speaks the model’s language by exporting FEN and SAN.

  1. Suppose the Anthropic API is down for an hour. Trace what a user can and cannot do in the app during that hour. Which layer is affected?

Hint: Only the AI layer makes network calls. The engine and UI keep working, so the user can still play full games against the local computer; they just lose commentary, the AI evaluation, and the Elo estimate. This is graceful degradation, and it is a direct consequence of the layering.

Which group of functions makes the HTTP requests to a language-model provider?

Answer

callTextModel, getAICommentary, getAIEvaluation

Common Misconception: “The AI plays the chess.” It does not. The computer opponent is the local minimax search in the engine layer: pure code, no network. The language model only comments on moves and estimates numbers. You could delete every AI function and still have a working chess game. Conflating “the program that plays” with “the model that talks” is the first confusion to clear up.


Part II: The Provider-Agnostic AI Layer

This is the heart of the activity. You will see how one function, callTextModel, sends the same prompt to Anthropic, OpenAI, or a local Open WebUI / Ollama server, differing only in URL, headers, and where the reply text sits in the response. This is the browser-JavaScript version of the provider portability you built in Python in the REST activity.

3. One Function, Three Providers

Everywhere else in the app, code that needs the model calls a single function:

const text = await callTextModel(providerConfig, prompt, { maxTokens: 350, temperature: 0.2 });

providerConfig is a plain object assembled from the UI (which provider is selected, plus the relevant key/URL/model). callTextModel reads providerConfig.provider and branches. Its skeleton is:

Code Cell

async function callTextModel(providerConfig, prompt, { maxTokens = 500, temperature = 0.2 } = {}) {
  const provider = providerConfig.provider;
  if (provider === "anthropic") { /* POST to api.anthropic.com/v1/messages */ }
  if (provider === "openai")    { /* POST to api.openai.com/v1/chat/completions */ }
  if (provider === "openwebui") { /* POST to <baseUrl>/api/chat/completions */ }
  throw new Error("Unsupported AI provider.");
}

Notice the shape: the rest of the app never knows or cares which provider is active. Adding a fourth provider later means adding one more branch here and nothing else changes. That is the whole point of a dispatcher.

4. Anthropic: POST https://api.anthropic.com/v1/messages

Anthropic’s Messages API uses its own endpoint and header names. Here is the branch, annotated:

Code Cell

if (provider === "anthropic") {
  if (!providerConfig.anthropicKey) throw new Error("Enter an Anthropic API key.");
  const model = providerConfig.anthropicModel || DEFAULT_MODELS.anthropic;

  const res = await fetch("https://api.anthropic.com/v1/messages", {
    method: "POST",
    headers: {
      "content-type": "application/json",
      "x-api-key": providerConfig.anthropicKey,          // <-- auth header (NOT "Authorization")
      "anthropic-version": "2023-06-01",                 // <-- required API version pin
      "anthropic-dangerous-direct-browser-access": "true" // <-- opt-in to call from a browser (see Part IV)
    },
    body: JSON.stringify({
      model,
      max_tokens: maxTokens,
      temperature,
      messages: [{ role: "user", content: prompt }],
    }),
  });

  const data = await res.json().catch(async () => ({ raw: await res.text() }));
  if (!res.ok || data?.error) throw new Error(data?.error?.message || data?.raw || res.statusText);

  // Anthropic returns a "content" array of blocks; join their text.
  return data?.content?.map(block => block?.text || "").join("").trim() || "";
}

Three things to note: the auth header is x-api-key (not Authorization: Bearer); a version header is mandatory; and the reply text lives at data.content[].text, an array of content blocks, not a single string.

5. OpenAI: POST https://api.openai.com/v1/chat/completions

The OpenAI branch uses the same /v1/chat/completions shape you already know from the REST activity:

Code Cell

if (provider === "openai") {
  if (!providerConfig.openaiKey) throw new Error("Enter an OpenAI API key.");
  const model = providerConfig.openaiModel || DEFAULT_MODELS.openai;

  const res = await fetch("https://api.openai.com/v1/chat/completions", {
    method: "POST",
    headers: {
      "content-type": "application/json",
      "authorization": `Bearer ${providerConfig.openaiKey}`, // <-- Bearer-token auth
    },
    body: JSON.stringify({
      model,
      messages: [{ role: "user", content: prompt }],
      max_completion_tokens: maxTokens,
      temperature,
    }),
  });

  const data = await res.json().catch(async () => ({ raw: await res.text() }));
  if (!res.ok || data?.error) throw new Error(data?.error?.message || data?.raw || res.statusText);

  return extractOpenAIText(data).trim(); // reads choices[0].message.content, defensively
}

The reply lives at choices[0].message.content. The app wraps that access in a small helper, extractOpenAIText, that tolerates a few response shapes instead of crashing on an unexpected one:

Code Cell

function extractOpenAIText(data) {
  if (typeof data?.output_text === "string" && data.output_text) return data.output_text;
  const msg = data?.choices?.[0]?.message;
  if (typeof msg?.content === "string") return msg.content;
  if (Array.isArray(msg?.content)) return msg.content.map(part => part?.text || part?.content || "").join("");
  return "";
}

6. Open WebUI / Local: POST <baseUrl>/api/chat/completions

The third branch targets a local, OpenAI-compatible server (Open WebUI in front of Ollama, or Ollama directly). Because it speaks the same protocol as OpenAI, the response parsing is identical; only the URL changes, and the key is optional:

Code Cell

if (provider === "openwebui") {
  const baseUrl = normalizeBaseUrl(providerConfig.openWebUIUrl); // trims trailing slashes
  if (!baseUrl) throw new Error("Enter an Open WebUI base URL.");
  if (!providerConfig.openWebUIModel) throw new Error("Choose an Open WebUI model.");

  const res = await fetch(baseUrl + "/api/chat/completions", {
    method: "POST",
    headers: {
      "content-type": "application/json",
      "authorization": providerConfig.openWebUIKey ? `Bearer ${providerConfig.openWebUIKey}` : "",
    },
    body: JSON.stringify({
      model: providerConfig.openWebUIModel,
      messages: [{ role: "user", content: prompt }],
      max_tokens: maxTokens,
      temperature,
    }),
  });

  const data = await res.json().catch(async () => ({ raw: await res.text() }));
  if (!res.ok || data?.error) throw new Error(data?.error?.message || data?.raw || res.statusText);

  return extractOpenAIText(data).trim(); // same choices[0].message.content path as OpenAI
}

The app can even discover which models a local server has, via a GET /api/models call in fetchOpenWebUIModels, the same “list models before you use one” pattern from the REST activity’s /v1/models.

Here is the whole idea, reduced to a runnable Python cell you can execute against your own local Ollama right now. It is the same request the browser makes (same endpoint shape, same messages array, same choices[0].message.content parse), just in Python:

Code Cell

Runs on your machine, not here. This cell talks to the Ollama server on your own laptop at localhost:11434, which a web page has no route to. Copy it into your course container and run it there.

import requests

def chat_completion(base_url, model, messages, api_key="ollama", temperature=0.2):
    """One request that works against ANY OpenAI-compatible server (Ollama, Open WebUI, cloud)."""
    endpoint = f"{base_url.rstrip('/')}/chat/completions"
    headers = {"Content-Type": "application/json", "Authorization": f"Bearer {api_key}"}
    payload = {"model": model, "messages": messages, "stream": False, "temperature": temperature}
    try:
        r = requests.post(endpoint, json=payload, headers=headers, timeout=120)
        r.raise_for_status()
        return r.json()["choices"][0]["message"]["content"]   # <-- the shared response path
    except Exception as e:
        import traceback; traceback.print_exc()
        return None

reply = chat_completion(
    base_url="http://localhost:11434/v1",   # local Ollama; swap to point anywhere
    model="llama3.2",
    messages=[{"role": "user", "content": "In one sentence, what does a chess coach do?"}],
)
print(reply)

Same Prompt, Three Response Shapes

The request bodies are nearly identical; the auth and the reply location differ. That table is the entire practical difference a developer must remember:

  Anthropic OpenAI Open WebUI / Ollama
URL api.anthropic.com/v1/messages api.openai.com/v1/chat/completions <baseUrl>/api/chat/completions
Auth header x-api-key: <key> + anthropic-version Authorization: Bearer <key> Authorization: Bearer <key> (optional)
Prompt goes in messages: [{role, content}] messages: [{role, content}] messages: [{role, content}]
Reply text is at data.content[].text data.choices[0].message.content data.choices[0].message.content
Needs a paid key? Yes Yes No (local model)

Questions to Work Through

  1. A student copies the OpenAI branch, changes the URL to Anthropic’s, and keeps Authorization: Bearer <key> and data.choices[0].message.content. Predict the two distinct failures they will hit and how each would appear at runtime.

Hint: (1) Auth: Anthropic ignores Authorization and wants x-api-key plus anthropic-version, so the request is rejected with an auth/version error before any reply. (2) Parsing: even with a valid reply, choices does not exist on an Anthropic response (the text is under content[].text), so choices[0] throws. Same lesson as the REST activity’s “empty response” bug, one layer up.

  1. The Open WebUI branch and the OpenAI branch parse the response with the same extractOpenAIText helper. What property of Open WebUI makes that reuse correct, and what would you have to change to add a provider that does not share it?

Hint: Open WebUI is OpenAI-compatible; it returns the same choices[0].message.content structure. A non-compatible provider (like Anthropic) needs its own parse branch, which is exactly why the Anthropic branch has its own content[].text line instead of calling extractOpenAIText.

  1. callTextModel throws a specific Error (e.g. “Enter an Anthropic API key.”) before it ever calls fetch when the key is missing. Why check first instead of letting the provider return a 401?

    Hint: A local guard gives a clear, instant, actionable message and avoids a pointless network round-trip (and a confusing provider-specific error body). Validate what you can locally; only spend a network call on things only the server can decide.

In a response from POST https://api.anthropic.com/v1/messages, where is the model’s reply text?

Answer

data.content[0].text (an array of content blocks)

Common Misconception: “If it’s the same prompt, it’s the same response object.” No. Providers agree on very little beyond “send messages, get a completion.” The request can look almost identical while the response shape and the auth headers differ. Write one small parse function per response family (OpenAI-style, Anthropic-style) and route to the right one; never assume choices[0] exists.


Part III: Prompt Engineering and Structured Output for Coaching

Now that the pipe exists, what do we push through it? Two different jobs: prose commentary (free text a human reads) and structured values (JSON the program uses to drive a meter or a badge). They demand different prompting.

7. The Coach Commentary Prompt

getAICommentary builds the prompt that produces the sentence-or-two after each move. Read how much context it assembles, and the guardrails it sets:

Code Cell

async function getAICommentary(providerConfig, beforeFen, afterFen, san, pgnSoFar, eloEst) {
  const prompt = `You are an expert chess coach. The player is White (estimated ~${eloEst} Elo).

You are analyzing the move that was just played, in the context of the entire game so far.

Full PGN so far (ending with the move being analyzed):
${pgnSoFar || "(opening)"}

Position BEFORE the move (FEN):
${beforeFen}

Move just played:
${san}

Position AFTER the move (FEN):
${afterFen}

Analyze ONLY the move just played, but do so in the context of the whole game so far.
Judge whether it is good, inaccurate, a mistake, or a blunder.
If it is suboptimal, name a better move for this turn and explain why.
Do NOT mention, predict, recommend, or hint at any future moves, replies, or continuations.
Keep the answer to 2-4 sentences.
Be encouraging but honest.`;

  return callTextModel(providerConfig, prompt, { maxTokens: 350, temperature: 0.2 });
}

Design decisions worth copying:

You can prove the prompt matters with a runnable cell, the same coach prompt, against your local model:

Code Cell

Runs on your machine, not here. This cell talks to the Ollama server on your own laptop at localhost:11434, which a web page has no route to. Copy it into your course container and run it there.


# Reuses chat_completion(...) from Part II.
fen = "rnbqkbnr/pppp1ppp/8/4p3/4P3/8/PPPP1PPP/RNBQKBNR w KQkq - 0 2"  # after 1.e4 e5
coach_prompt = f"""You are an expert chess coach. Analyze ONLY the move just played, in 2-3 sentences.
Position after the move (FEN): {fen}
Move just played: e5
Judge it as good, inaccurate, a mistake, or a blunder. Do not predict future moves."""

print(chat_completion(
    base_url="http://localhost:11434/v1", model="llama3.2",
    messages=[{"role": "user", "content": coach_prompt}],
))

8. Structured Output: Asking for JSON You Can Use

The evaluation bar and the Elo meter cannot use a paragraph; they need a number. So those prompts demand JSON and the code parses it. Here is the evaluation call, start to finish:

Code Cell

async function getAIEvaluation(providerConfig, fen) {
  const prompt = `You are a chess engine. Given this FEN, estimate the evaluation from White's` +
    ` perspective as a single number (positive = White advantage, negative = Black advantage).` +
    ` Respond ONLY with a JSON object like {"eval": 0.5} where the number is in pawns. FEN: ${fen}`;

  const text = await callTextModel(providerConfig, prompt, { maxTokens: 120, temperature: 0 });

  // Models often wrap JSON in ```json fences or add stray words. Strip fences, then parse safely.
  const parsed = safeJsonParse(text.replace(/```json|```/g, "").trim(), {});
  return typeof parsed.eval === "number" ? parsed.eval : 0;   // fall back to 0 if anything is off
}

function safeJsonParse(text, fallback) {
  try { return JSON.parse(text); } catch { return fallback; }
}

Three defenses stacked together make this robust:

  1. Ask precisely: “Respond ONLY with a JSON object like {"eval": 0.5}” and temperature: 0.
  2. Clean the text: strip ```json fences the model may add.
  3. Never trust the parse: safeJsonParse returns a fallback instead of throwing, and the code then checks that parsed.eval is actually a number before using it.

The Elo estimator (getSideAIElo) does the same for {"elo": 1200, "label": "Intermediate"}, with parsed.elo || 1200 and parsed.label || "Unknown" as fallbacks.

9. Async Orchestration and the Stale-Closure Trap

After you move, the app wants two independent AI answers: commentary and a fresh Elo estimate. They don’t depend on each other, so it fires them together with Promise.all and waits once:

Code Cell

const [c, elo] = await Promise.all([
  getAICommentary(currentProviderConfig, beforeFen, afterFen, san, pgnSoFar, eloEstimate.elo),
  getAIElo(currentProviderConfig, afterFen, newSans),
]);
setCommentary(c);
setEloEstimate(elo);

One subtlety specific to long-lived UI code: the move handler is an async function that may still be awaiting a slow API call when you change the provider dropdown. If it read the provider from the React render that created it, it could use stale settings. The app avoids this by reading the current config from a ref at call time:

Code Cell

const currentProviderConfig = providerConfigRef.current; // always the latest, even mid-await
const currentAIEnabled = aiEnabledRef.current;
// ...a useEffect keeps providerConfigRef.current in sync whenever the settings change.

Tracing One Move Through the AI Layer

You are White. You play Nf3. Walk the sequence the app performs (see executeMove).

Step What happens Layer
1 moveToSAN names the move "Nf3"; applyMove produces the new state Engine
2 UI updates instantly: piece moves, move list appends React UI
3 If AI is enabled: Promise.all fires getAICommentary and getAIElo AI layer
4 Each builds a prompt (FEN + SAN + PGN), calls callTextModel -> fetch AI layer
5 Replies parsed (prose as text; Elo as JSON) and shown AI layer
6 The local minimax picks Black’s reply and the board updates Engine

Questions to Work Through

  1. Steps 1-2 (engine + UI) finish in well under a millisecond; steps 3-5 (AI) can take several seconds. The app updates the board before awaiting the AI. Why is that ordering a deliberate UX decision, not an accident?

Hint: The move is already legal and known; there is no reason to make the human wait on a network call to see their own move. Render the certain, cheap result immediately; stream in the slow, uncertain AI result when it arrives. Never block a deterministic UI update on a probabilistic network call.

  1. getAIEvaluation strips ```json fences before calling JSON.parse. Give a concrete model output that would make JSON.parse throw without that strip, and explain why safeJsonParse still keeps the app alive even if the strip missed something.

    Hint: A reply like ` json\n{"eval": 0.3}\n ` is not valid JSON because of the fence lines; JSON.parse throws on the backticks. safeJsonParse wraps the parse in try/except and returns the fallback {}, so the app shows a neutral eval instead of crashing.

  2. Suppose the Elo call succeeds but returns {"label": "Intermediate"} with no elo field. What number ends up on the meter, and which line of code decided that?

Hint: parsed.elo || 1200 supplies 1200 when elo is missing/falsy. Defensive defaults on every field mean a partially-malformed structured response degrades to something sensible instead of undefined reaching the UI.

Why does getAIEvaluation call safeJsonParse instead of JSON.parse directly?

Answer

So a malformed or fenced model reply returns a fallback instead of throwing and breaking the render

Common Misconception: “If I ask for JSON, I get JSON.” Language models are usually obedient but never guaranteed. They add prose, wrap output in code fences, or drop a field. Treat every structured reply as untrusted input: constrain the prompt, strip known wrappers, parse defensively, and default every field. Robust structured output is 20% prompt and 80% parsing discipline.


Part IV: Securing Your API Keys

This part is not optional polish; it is the difference between a safe project and a leaked credential that runs up someone else’s bill. The rules are simple; the failures are expensive.

10. The Golden Rule: Never Commit or Hardcode a Key

A key like sk-REPLACE_ME... is a password. The two ways students most often leak one:

How the Chess AI Coach avoids both: the key is only ever typed into a password field and held in React state (useState) for the session. It is never written to disk, never hardcoded, and there is nothing to commit. Close the tab and it’s gone. In a project with a build step, the equivalent discipline is: read keys from environment variables and add your .env file to .gitignore so it can never be committed.

11. The Client-Side Exposure Problem

Here is the uncomfortable truth about the app’s Anthropic branch. It calls api.anthropic.com directly from the browser, and Anthropic makes you opt in with a header literally named:

anthropic-dangerous-direct-browser-access: true

The word “dangerous” is doing real work. When the browser makes the call, the key travels from the user’s browser and is visible in that browser’s DevTools -> Network tab on every request. Think about who can see it in each situation:

type="password" on the input only hides the characters from someone looking over your shoulder. It does nothing to hide the key from the network layer or from other scripts on the page.

12. The Production Pattern: A Backend Proxy

For any app real users will touch, the key belongs on a server you control, never in the browser:

   Browser (no key)            Your backend (holds key)          Provider
  +---------------+   POST    +----------------------+  POST   +----------+
  |  chess UI     | --------> |  /api/coach          | ------> | Anthropic|
  |  fetch("/api")|           |  key = os.environ[...] |         | / OpenAI |
  `---------------+ <-------- `----------------------+ <------ `----------+
        reply                        reply

The browser calls your endpoint. Your server reads the key from an environment variable and adds it to the provider request. The secret never leaves your server; users never see it; and you can add rate limits and logging in one place. A minimal proxy is only a few lines:

Code Cell

Runs on your machine, not here. This cell starts a server and binds a port, which a web page cannot do. Copy it into your course container and run it there.


# minimal_proxy.py  - run with a real key in the environment, e.g.

#   export PROVIDER_API_KEY="sk-...."   (never commit this)
import os, requests
from flask import Flask, request, jsonify

app = Flask(__name__)
API_KEY = os.environ["PROVIDER_API_KEY"]  # <-- from the environment, NOT the source code

@app.post("/api/coach")
def coach():
    user_prompt = request.get_json()["prompt"]
    r = requests.post(
        "https://api.openai.com/v1/chat/completions",
        headers={"Authorization": f"Bearer {API_KEY}"},   # key added server-side
        json={"model": "gpt-4.1-mini",
              "messages": [{"role": "user", "content": user_prompt}]},
        timeout=120,
    )
    data = r.json()
    return jsonify({"text": data["choices"][0]["message"]["content"]})

The local-model escape hatch sidesteps the whole problem: point the app at Open WebUI / Ollama and there is no cloud key to leak. That is a big reason this course leans on local models, and why the app supports them as a first-class provider.


Where Should the Key Live?

Three deployment scenarios. For each, decide where the key should live.

Scenario Who uses it Safe place for the key
A. You experiment on your own laptop Only you In the browser session (typed into the field), fine
B. A shared classroom instance for 20 students Many people On a backend, or use a keyless local model
C. A public website anyone can visit The whole internet On a backend proxy; never in the browser

Questions to Work Through

  1. In scenario C, a teammate suggests “we’ll just obfuscate the key by base64-encoding it in the JavaScript so nobody can read it.” Explain precisely why this does not work.
> *Hint: Anything the browser can decode to make the request, an attacker can also decode; the browser must send the real key over the wire, so it appears decoded in the Network tab regardless of how it was stored in the source. Obfuscation is not encryption; the secret still reaches the client. The only fix is to not send the secret to the client at all.*
  1. Scenario B has two safe options. Compare them: what does the backend-proxy option buy you that the keyless-local-model option does not, and vice versa?
> *Hint: The proxy lets you use powerful cloud models while centralizing the key, rate limits, and logging, at the cost of running a server and paying per call. The local model needs no key and no per-call cost and keeps data on-premises, at the cost of hardware and (often) lower capability. Different tradeoffs, both avoid a client-side secret.*
  1. The app stores the key in useState, not in localStorage. Why is that a safer default for a browser app, even for personal use?
> *Hint: `useState` lives only in memory for the tab's lifetime, so the key vanishes when the tab closes and never persists to disk where other scripts or a shared machine's next user could read it. `localStorage` would survive restarts and be readable by any script on the origin, more exposure for no real benefit here.*

In a safe public deployment, where does the provider API key live?

Answer

On a backend server you control, read from an environment variable

Common Misconception:type='password' protects the key.” It only masks the characters on screen. The key is still in memory, still sent over the network in plain view of the Network tab, and still readable by any script on the page. Masking ≠ protecting. The real protections are: keep the key off the client entirely (backend proxy) or use a provider that needs no key (local model).


Part V: Wiring the AI into the User Interface

The final layer is glue: React state that holds the provider settings (never keys in code), and the move handler that decides whether and when to call the model.

13. Provider Config State and executeMove

The provider settings are ordinary React state (a dropdown value plus the relevant fields), assembled into providerConfig and handed to the AI functions. The move handler branches on whether AI is configured:

Code Cell

// Is any provider actually ready to use?
const aiEnabled =
  (provider === "anthropic" && !!anthropicKey.trim()) ||
  (provider === "openai"    && !!openaiKey.trim()) ||
  (provider === "openwebui" && !!openWebUIUrl.trim() && !!openWebUIModel);

// Inside executeMove, after the move is already applied and drawn:
if (currentAIEnabled) {
  setCommentaryLoading(true);
  try {
    const [c, elo] = await Promise.all([ getAICommentary(...), getAIElo(...) ]);
    setCommentary(c);
    setEloEstimate(elo);
  } catch (e) {
    setCommentary(`AI request failed: ${e.message}`); // errors become visible, not silent
  } finally {
    setCommentaryLoading(false);
  }
} else {
  setCommentary(`You played ${san}. Configure a provider for AI analysis.`);
}

Notice the try/catch/finally: a failed API call turns into a visible message and the loading spinner always clears. Silent failures are the worst failures in AI features, because the user can’t tell “the model is thinking” from “the model is broken.”

14. Graceful Degradation, One More Time

Because aiEnabled gates every model call, the app has a complete fallback path:

This is the template for adding AI to any existing app: make the app fully work without the model first, then layer the model on as an enhancement that fails safe.


Exercises

  1. Add a fourth provider.
  1. Add a structured-output feature.
  1. Swap the coach’s persona.
  1. Move the key server-side.

Reflection Prompt

Personal: Before reading this app, did “the AI analyzed my move” feel like magic? Now that you have seen it is an HTTP POST with a carefully worded prompt and a defensive JSON parse, has your sense of what these features are changed? What still feels non-obvious?

Technical: In your notebook, design the configuration and secret-handling for a version of this app that your whole class could use at once. Where does the key live? Which provider(s) do you support and why? Sketch the request path from a student’s browser to the model and back, and mark every place a secret must not appear.

Societal: The app estimates a player’s Elo from their moves and shows it back to them in real time. What are the risks of software that continuously scores a person’s skill and reports it? Who might be discouraged or misjudged by a wrong estimate, and what responsibilities does a developer have when a model’s confident-looking number is actually a rough guess?


Where This Goes Next

You now have the full pattern for adding a language model to real software: isolate the AI layer, dispatch across providers, engineer prompts for prose and for structured JSON, parse defensively, and keep secrets off the client. In the Build Your Own AI Coach lab you will apply exactly this pattern to a domain of your choosing (a simpler game, a writing tutor, a code reviewer), reusing the provider-agnostic call, the structured-output discipline, and the key-security rules you practiced here.


Further Reading