November 2023. OpenAI launched the function calling API with stable support, and without any formal announcement, the trade of “prompt engineer” began its decline. Not because prompts stopped mattering, but because they stopped being what mattered most.
At Soamee, we’ve been building products with LLMs for years. We went through the magic-prompt phase, through the blind faith that perfect message wording was the secret to success. And we learned, sometimes painfully, that this faith was misplaced.
This article is my direct take on what happened, why prompt engineering as an isolated discipline is dead, and what truly moves the needle in AI products in 2026.
The Era of Magic Prompts (2022-2024)
When GPT-3 went mainstream and GPT-4 popularized it, the tech world split into two groups: skeptics who thought it was all hype, and enthusiasts who believed they’d found the ultimate oracle. The enthusiasts were right about one thing: LLMs were genuinely powerful. But they drew the wrong conclusion about why they worked.
The dominant narrative was: “The key is knowing how to talk to the model.” Prompt engineering courses appeared at hundreds of dollars a pop. Repositories with thousands of stars filled up with “definitive prompts.” LinkedIn overflowed with profiles calling themselves “Prompt Engineer” as if it were the career of the future.
The problem wasn’t that prompts didn’t matter. The problem was the conclusion being drawn: that the difference between an excellent AI product and a mediocre one lay in having found the right phrasing. That there were magic prompts waiting to be discovered.
This idea was wrong. And it took the market between 18 and 24 months to realize it.
What Killed Prompt Engineering
1. Structured outputs made format prompts obsolete
For years, a significant part of “prompt engineering” work was getting the model to respond in the correct format. You’d ask for JSON and sometimes get JSON, sometimes text containing JSON, sometimes something that looked like JSON but wasn’t. Entire libraries existed dedicated to parsing and correcting model outputs.
Today, OpenAI, Anthropic, and Google natively support structured outputs: you define a JSON schema and the model guarantees its response conforms to that schema. No magic, no perfect wording. A technical contract.
# Before: praying the model would respond in JSON
response = client.messages.create(
model="claude-3-5-sonnet",
messages=[{"role": "user", "content": "Extract the name and email. Respond ONLY in JSON with fields 'name' and 'email'. Do NOT include additional text."}]
)
# Result: unpredictable
# Now: define the contract
class ExtractedContact(BaseModel):
name: str
email: str
response = instructor.from_anthropic(client).messages.create(
model="claude-3-5-sonnet",
response_model=ExtractedContact,
messages=[{"role": "user", "content": "Extract the contact from the text"}]
)
# Result: guaranteed
The format prompt is dead. Engineering killed it.
2. Tool use made action prompts obsolete
The second major pillar of prompt engineering was getting the model to make the right decisions: “If the user asks X, do Y. If they ask Z, do W.” Prompts with conditionals, decision trees written in natural language.
Function calling replaced this elegantly. Instead of trying to get the model to simulate business logic through text, you give it real tools and let it decide when to use them. The model is good at high-level decision-making; code is good at executing actions precisely.
tools = [
{
"name": "search_order",
"description": "Search for an order by ID or customer email",
"input_schema": {
"type": "object",
"properties": {
"query": {"type": "string"},
"type": {"enum": ["id", "email"]}
}
}
},
{
"name": "update_status",
"description": "Update the status of an order",
"input_schema": {
"type": "object",
"properties": {
"order_id": {"type": "string"},
"new_status": {"enum": ["processing", "shipped", "cancelled"]}
}
}
}
]
You don’t need to write “If the user mentions an order number, search for it first before responding.” The model infers it from the tool context. What was once a paragraph of instructions in the prompt is now a function definition.
3. Context engineering matters more than prompt wording
Here’s the insight that has cost our industry the most to internalize: what you put in the context matters far more than how you word the question.
You can have the most elegant prompt in the world, but if the model doesn’t have access to the relevant information, it will hallucinate or give generic answers. And you can have a basic, almost telegraphic prompt, but if you give it the right information, the model will perform well.
Context engineering is the discipline of designing what reaches the model before it responds:
- Well-implemented RAG: Not just embeddings and cosine similarity. Selecting relevant chunks, reranking, filtering by metadata, sufficient context without exceeding limits.
- Dynamic few-shot examples: Not static examples in the prompt. Examples dynamically selected based on the current query.
- Conversational context management: Which previous turns to keep, which to compress, which to discard.
- Structured data vs. text: Sometimes it’s better to give the model a well-formatted table than a paragraph explaining the data.
In our projects, 80% of quality improvements have come from improving the context, not from rewriting the prompt.
4. Evaluation replaced gut feeling
Classic prompt engineering was fundamentally vibes-driven. You’d try a variation, it seemed better, you’d ship it. There was no rigorous way to measure whether it was actually better across all cases or just the ones you’d tested manually.
The field’s maturity has brought evals: sets of test cases with explicit evaluation criteria, often evaluated by another LLM. Frameworks like RAGAS, LangSmith, Braintrust, or PromptFoo allow this to be done systematically.
When you have evals, you can do what you did with code: measure before changing, compare versions, detect regressions. The prompt stops being a handcrafted artifact no one touches for fear of breaking something and becomes versioned code with tests.
This changes the nature of the work. It’s not intuition—it’s engineering.
What Replaced It: LLM Engineering
What’s emerging isn’t evolved prompt engineering. It’s an entirely different discipline that borrows the name from software engineering because that’s essentially what it is.
LLM engineering treats the language model as just another system component: with defined interfaces, data contracts, automated tests, observability, and continuous deployment. The prompt is a versioned configuration file in git, not text saved in an environment variable that only the team’s expert touches.
The pillars of this discipline:
System prompts as software specifications: Versioned in git, reviewed in pull requests, with regression tests. If someone changes the system prompt, the CI pipeline runs the evals and blocks the merge if quality drops.
Structured outputs as API contracts: The model is an internal service with a defined interface. What it returns has a schema. The rest of the system trusts that schema, not parsing free text.
Tool use as agent architecture: Instead of trying to make the model do everything, you give it specialized tools and let it orchestrate. Business logic lives in the tools (testable code), not in the prompt.
Observability and traceability: Every model call is logged with its full context, tokens used, latency, result. When something fails in production, there’s a trace that tells you exactly what happened.
Evals as CI/CD for prompts: Before deploying any change to the system prompt or RAG architecture, you run the eval suite. If the score drops, it doesn’t deploy.
Where Prompting Still Matters
It would be dishonest to say prompts don’t matter at all. They do. But in specific contexts:
Creative tasks: When you want the model to generate text with a particular style, the wording of the prompt remains an art. “Write in the style of an Economist journalist covering technology” produces different results from “Write an article about technology.” Here linguistic nuance matters.
Rapid prototyping: When you’re validating an idea in a couple of hours, it doesn’t make sense to build all the eval infrastructure and structured outputs. A well-written prompt in Cursor or the Anthropic console is perfectly valid for exploring feasibility.
Edge cases and reasoning behavior: For tasks requiring chain-of-thought reasoning, how you structure the problem in the prompt still has impact. “Think step by step” and its variations aren’t magic, but they are instructions that affect the reasoning process.
Models without tool use or structured outputs: If for some reason you’re working with smaller or locally deployed models that don’t support these capabilities, classic prompting remains relevant.
What’s dead isn’t the prompt itself. What’s dead is the idea that the prompt is the central artifact around which everything else revolves.
What Your Team Should Learn
If you’re building products with LLMs in 2026, these are the skills that actually matter:
1. JSON schema design: Knowing how to define clear data structures for model outputs. Pydantic, Zod, JSON Schema. This is more important than knowing how to write prompts.
2. Agent architectures: Understanding how to design systems where the LLM orchestrates specialized tools. Patterns like ReAct, Plan-and-Execute, and when to use each.
3. Production RAG: Not the 5-minute tutorial with ChromaDB and basic embeddings. Real RAG includes strategic chunking, reranking, relevance evaluation, and quality monitoring.
4. Eval design and execution: Knowing how to define what “good” means for your use case, build representative test sets, and run evaluations systematically.
5. LLM observability: Tracing, logging, cost and latency monitoring. Tools like LangSmith, Langfuse, or Helicone. You can’t improve what you can’t measure.
6. Context management: Understanding context limits, when and how to compress, when to use external memory vs. in-context context.
Prompts are part of all this. But they’re one component among many, not the center.
Conclusion: Build Systems, Not Prompts
The difference between a team that uses AI productively and one that doesn’t isn’t who knows how to write better prompts. It’s who has built the right infrastructure around the model: the data contracts, the tools, the evaluation pipeline, the observability.
The prompt engineer who only knows how to write messages is the equivalent of a frontend developer who only knows how to copy code from Stack Overflow without understanding why it works. Useful for one-off tasks, insufficient for building quality products.
What the market demands now are engineers who understand LLMs as system components: their capabilities, their limitations, and how to integrate them robustly into real software architectures.
At Soamee, we’ve been working exactly this way for a while. We don’t sell magic prompts or ChatGPT workshops. We build artificial intelligence products with the same engineering discipline we’d apply to any other software system: defined contracts, tests, observability, and continuous deployment.
If your company is evaluating how to integrate AI seriously, beyond the initial prototype, let’s talk. Book a free consultation and we’ll discuss architecture, not prompts.