Lesson 4 of 8 · Build AI Apps with Real APIs
In this lesson. Hide wait time and avoid paying twice for the same prompt.
What you will learn
- Streaming versus waiting for the full reply
- What is safe to cache
- A back-of-the-envelope cost estimate
Walkthrough
Users feel latency before they feel intelligence. You will stream tokens when it helps, show a status for slow tools, and cache repeated prompts. You will also estimate cost per 1,000 requests so a “clever” extra call is a product decision, not an accident.
Work through the ideas in order. After each point, pause and connect it to a task you already do — a document, a workflow, or a feature you own. The goal of Build AI Apps with Real APIs is usable skill, not a pile of notes.
If something is unclear, rewrite it in your own words before you continue. Teaching the step back to yourself is the fastest way to see gaps.
Practice
Storyboard one AI screen: what the user sees at 0s, 2s, and 8s, and which calls you would cache.
Keep the first attempt small. A finished example you can reuse beats a perfect plan you never run.
Check your understanding
- Can you explain the goal of this lesson in one sentence to a teammate?
- Where would you apply “Streaming versus waiting for the full reply” in your own work this week?
- What would you change on a second pass of the practice?
Next. Continue to the following lesson when the practice has a real artifact, even a rough one.
