Insights
This a go back link
Artificial Intelligence
,
ChatGPT
,
OpenAI
,
GitHub Copilot
,
Generative AI
,

GPT-5.6 Luna in GitHub Copilot: Cost Optimization, Subagents, and Context Management

GPT-5.6 Luna in GitHub Copilot: Cost Optimization, Subagents, and Context Management
24.9.2026

GPT-5.6 Luna breaks the old rule that cheap models are only good for basic tasks. With MAX reasoning and 1M context, it handles research, implementation, and subagent work at a fraction of frontier-model cost. Luna is the new Swiss Army knife for daily development.

Lightweight models used to mean boilerplate and quick edits. Luna is good enough for serious work and cheap enough to use by default. Save frontier models for the hardest decisions.

GPT-5.6 Luna Benchmarks: Redefining Lightweight AI Models

GPT-5.6 Luna is a lightweight OpenAI model available in GitHub Copilot.

Artificial Analysis scores Luna at 52; it performs close to Sonnet 5 and well above GPT-5.4 mini. The measured cost per benchmark task is about $0.05 for Luna versus $1.72 for Sonnet 5.

OpenAI on x.com

The benchmark used Luna with MAX reasoning.

For the direct OpenAI API, standard GPT-5.6 Luna pricing per one million tokens is:

GitHub Copilot doubles Luna's rates above 200K input tokens. It remains far cheaper than frontier models. I recommend using it with thinking effort on MAX.

Context Economics: Why Long GitHub Copilot Chats Cost More

The model does not receive only your latest message. Each turn can resend system instructions, tool definitions, attached files, earlier messages, and tool output.

That means a 100K-token chat used for five more turns can process much of the same 100K tokens five more times. Prompt caching reduces the price of repeated context, but the chat still grows toward its context limit.

Use /compact when the same task grows long. Start a new chat when the topic changes.

Use Cheap Subagents to Protect Frontier Context

A frontier parent model can delegate focused work to cheaper subagents. Good tasks include mapping a codebase, gathering documentation, comparing options, or running tests.

Each subagent gets a clean context and returns only its final result. Its research and tool output never enter the expensive parent's context.

Check the model before delegating. Subagents inherit the parent model by default, and every subagent processes its own initial prompt. Accidentally running several frontier subagents can cost more than doing the work in the parent chat.

With a cheap model, the subagent handles its calls cheaply and the frontier parent receives a short answer instead of pages of intermediate work.

The parent model decides whether to delegate and writes the subagent task. You can steer it in your prompt: e.g., "Use GPT-5.6 Luna subagents for online research."

Configuring Copilot Subagents in VS Code (Explore, Search, Implement)

GitHub Copilot in VS Code now uses several specialized workers:

  • Explore: Read-only subagent used by Plan to map the codebase.
  • Search: Experimental subagent that searches iteratively in an isolated context.
  • Implement: Plan's coding step, with a separate model setting.

To use a different model for this work:

  1. Enable agent/runSubagent in the chat tools picker.
  2. Enable github.copilot.chat.exploreAgent.enabled for Search and exploration.
  3. Set chat.exploreAgent.defaultModel to e.g., GPT-5.6 Luna.
  4. Ask the parent to delegate research or implementation to subagents.

GitHub Copilot Prompt Caching: How to Preserve Parent Cache

GitHub Copilot caches the unchanged prefix of a conversation. Cached tokens are typically billed at 10% of the normal input rate, so preserving that prefix matters most in long sessions.

GitHub states that Copilot caches expire after 24 hours of inactivity for OpenAI models and after 1 hour for most others. Treat these figures cautiously: provider defaults can be shorter, and Copilot does not expose a clear cache-expiration contract in its billing documentation. From experience, most model caches expire after 5 minutes.

This matters when a frontier parent waits several minutes for a subagent. If the parent cache expires, the next turn must rebuild its context, which can offset some of the savings from using a cheap subagent.

Use subagents primarily to keep research and tool output out of the parent context, not as a guaranteed cost reduction. Keep their tasks focused, use a cheaper model, and check Agent Debug Logs to see whether the parent still receives cache hits. For short, simple work, staying in the parent can be cheaper. Even a short implementation task may be cheaper on the same expensive model while the parent cache is still warm.

Inspect real sessions when you want to confirm whether the cache survived:

  1. Enable github.copilot.chat.agentDebugLog.fileLogging.enabled in VS Code Settings.
  2. Open the Chat view's ellipsis menu and select Show Agent Debug Logs.
  3. Select the session name in the breadcrumb to open its Summary view.
  4. Select Cache Explorer to see cache-hit rates and the first prompt difference that caused a miss.

Pro Tips for Optimizing GPT-5.6 Luna & Copilot Workflows

  • Always use Luna with MAX reasoning and 1M context. The small credit increase is worth it and remains far below frontier-model cost.
  • Use Luna for research and implementation. Escalate only when needed.
  • Put noisy research in read-only subagents.
  • Use /compact for long tasks and a new chat for new topics.
  • Check Cache Explorer instead of guessing about cache hits.
  • Keep cache expiry in mind. A short, focused conversation is often better than leaving long gaps between prompts and risking a cache miss.
  • For long-running tool calls, enable github.copilot.chat.agent.longToolCallCachePreservation.enabled to send periodic keep-alive probes that keep the server-side prompt cache warm. Use github.copilot.chat.agent.longToolCallCachePreservation.maxProbes to limit how many probes VS Code sends before giving up. Both settings are experimental as of now.

Update — September 22, 2026: GitHub has announced the gradual rollout of GPT-6 Sol and GPT-6 Luna in Copilot. GPT-6 Luna continues the focus on lightweight, cost-efficient assistance explored here with GPT-5.6 Luna, with input token rates cut in half and output token rates reduced by more than half compared with its predecessor at standard pricing. For teams evaluating the new model, the same considerations apply: task suitability, subagent delegation, and the impact of context and caching on overall costs.

Oskar Arden
Oskar Arden
Software Architect
A Software Architect and Fullstack Engineer specializing in CI/CD pipelines, cloud DevOps, and enterprise AI integration. He focuses on scaling modern application architecture, prompt caching strategies, and optimizing AI-driven development workflows.

More related topics

This is a a back to top button