Skip to content
SpecifyDocs

Models

Choosing a model

Recommendations, based on internal evaluations, for which model and reasoning effort to use for each kind of task.

If you're not sure which model to use, start with GPT-6 Luna, the default model. For writing documents, multi-step editing, or tasks that need to follow a skill, we recommend Claude Sonnet 4.6. For the hardest analysis and for verifying long documents, use Claude Opus 4.6.

Recommendations by task

What you want to doRecommended modelReasoning effortWhy
Search team docs, ask short questions, summarizeGPT-6 LunaOffThe least expensive default model. In the Haiku and Sonnet comparison, search- and lookup-focused tasks showed no difference between models.
Quick lookups with ClaudeClaude Haiku 4.5OffShowed a success rate similar to Sonnet for lookups, summaries, and search.
Write new documents or edit several places in an existing oneClaude Sonnet 4.6Low to MediumHigher success rate than Haiku on document editing tasks.
Write documents that follow a skill or plugin procedureClaude Sonnet 4.6Low to MediumSkill application tasks showed the largest quality gap.
Tasks that use connectors such as GitHub or SlackClaude Sonnet 4.6LowCompleted every connector scenario.
Design the structure of long documents, verify quality, handle difficult analysisClaude Opus 4.6Medium to HighThe deepest reasoning. It's the most expensive, so use it only in the conversations that need it.
Compare the same task with a model from another familyGPT-6 SolMediumA high-performance OpenAI alternative that costs less than Sonnet.

Internal evaluation results

Specify compared Claude Haiku 4.5 and Claude Sonnet 4.6 across 38 real workspace scenarios, including document editing, skill application, connectors, memory, RAG search, and multi-turn conversations.

MetricClaude Haiku 4.5Claude Sonnet 4.6
Task success rate66%78%
Grader score (out of 5)4.194.59
Cost per successful taskBaselineAbout 4.3×

Success rates by task type:

Task typeHaiku 4.5Sonnet 4.6
Skill application26%80%
Document editing66%77%
Connectors50%100%
RAG search83%83%
Multi-turn conversation83%83%
Avoiding cost traps80%80%
Memory60%60%

Conclusion: Tasks that follow a skill or edit documents in multiple steps need Sonnet. Search-, summary-, and lookup-focused tasks work well with a less expensive model.

This evaluation compared only Claude Haiku 4.5 and Sonnet 4.6 (July 2026). GPT models and Opus weren't measured under the same conditions. The recommendations above for GPT models and Opus are based on specifications such as price and context, and on Specify's internal defaults.

Specify's internal defaults

When Specify picks a model itself in the background, it follows the same principles.

TaskDefault model
Maintaining and updating documents, planningClaude Sonnet 4.6
Research, synthesizing results, analysisClaude Haiku 4.5
Designing and verifying document quality outlinesClaude Opus 4.6
Document automation (when no workspace setting exists)GPT-6 Luna

Choose a reasoning effort

  • Off: Search, summaries, and short answers. The fastest.
  • Low · Medium: Writing and editing documents, and questions that compare several sources.
  • High and above: Problems where the right answer is hard to find, such as design decisions and complex root cause analysis. Responses slow down and usage increases significantly.

Set the workspace default models

Automations and document generation use the models in workspace settings, separately from chat.

The Default AI Model and Doc Generation Model pickers in Workspace General settings
  • Default AI Model: If you have many lightweight automations such as weekly summaries, we recommend GPT-6 Luna. If you have many automations that edit documents, we recommend Claude Sonnet 4.6.
  • Doc Generation Model: If document quality matters, we recommend Claude Sonnet 4.6. If cost comes first, start with the default (GPT-6 Luna), then change it after reviewing the results.