Evaluating an LLM Is Different From Evaluating a RAG System

I recently read Hamel Husain and Shreya Shankar’s article about similarity metrics. Their argument is that metrics like ROUGE and BERTScore often miss the failures that matter in an LLM application. I agree. But I think we need to separate two things: evaluating what a model learns during training and evaluating whether a RAG application works. Both involve an LLM, but the questions are different. If you do not understand the question, choosing a metric will not help you much. ...

October 6, 2026 · 6 min · 1227 words · Necati Demir

I Tested Laya After the Jev Announcement. Then I Fine-Tuned It.

1. From Jev to Laya I recently read TypeSafe’s announcement of Jev. The idea is interesting: give a model some information and predefined questions, then get structured decisions with probabilities. A support workflow, for example, could use it to choose which team should handle a ticket. But a structured answer can still be wrong. I wanted to see how this kind of model performed on a concrete task. Then I found Laya, an open source decision model. Its author, Nandakishor M, claims he worked on this direction “a year ago.” His March 2025 SalesRLAgent paper documents earlier sales-prediction research. His write-up describes developing the general-purpose model afterward, so I would not say today’s Laya checkpoint was released a year ago. ...

September 21, 2026 · 9 min · 1867 words · Necati Demir

Claude Got Banned From Government Systems --- Here's the Lesson.

Imagine you’re a CTO at a company with a government contract. Your product runs on Claude from Anthropic. Monday morning, you get a call: Claude is banned from all government systems. Now what? Two Very Different Mornings If you built your AI layer as a plug-and-play abstraction, you’re annoyed, but you’re fine. You swap the provider, run your regression tests, push to staging, move to prod. That’s it. A day’s work. Maybe two. ...

April 7, 2026 · 2 min · 341 words · Necati Demir

Sam Altman's Worst Take: "Don't Learn to Code"

Sam Altman recently said something that caught my attention: “Don’t learn to code.” Instead, he says, learn high agency, soft skills, and idea generation. I think he’s half right—and half dangerously wrong. Yes, high agency, soft skills, and idea generation matter. They mattered before the LLM era, and they still matter now. No argument there. But telling people to stop learning to code? That is the most dangerous advice in tech right now. ...

March 6, 2026 · 3 min · 436 words · Necati Demir

3 Types of Developers. Only One Will Survive.

The Observation I recently watched a new graduate join a team. Fresh out of school. Zero real-world experience. Zero familiarity with the codebase or tech stack. That’s typical—every new hire starts there. What wasn’t typical: he went from “where is the repo?” to helping the team with non-trivial tasks in just a couple of weeks. Not months. Weeks. What Changed Twenty years ago, ramping up in a complex codebase with an unfamiliar tech stack took a minimum of a couple of months to become fully productive. That was normal. Nobody expected you to be fully functional in weeks. ...

January 5, 2026 · 3 min · 461 words · Necati Demir

Your Company's AI Strategy Is EMBARRASSING And Everyone Knows It

Your company just announced a new AI initiative. You hired a Head of AI. You’re “exploring use cases.” You have 56 meetings on your calendar about it. And your engineers? They’re laughing at you. Not with you. AT you. The worst part? You have no idea. After two decades in technology and software engineering, I’ve watched this pattern repeat itself across countless organizations. Let me show you exactly why your AI strategy is failing—and what you can do about it. ...

December 17, 2025 · 4 min · 665 words · Necati Demir

The AI vs Teen Driver Comparison Is Wrong (Here's What Everyone Forgets)

The Popular Comparison A teenager learns to drive in 10 hours, but an AI system needs millions of simulations and millions of hours of simulated data. This comparison appears frequently in AI discussions, but once you look closely, it’s not fair. Here’s why. Where This Example Comes From I recently saw this example in an interview with Ilya Sutskever, co-founder of OpenAI who later started his own AI startup with significant investment backing. ...

December 3, 2025 · 4 min · 729 words · Necati Demir

Stop Letting AI Write Your Code: The SCORE Framework for Responsible AI Development

Developers are using AI tools like Cursor, Claude Code, GitHub Copilot, and Codex every day. But most are doing it wrong. Here’s what you need to understand: Your job is not to let the AI write code. Your job is to stay in control of the change. The Problem Yes, AI coding tools are really good now. You can literally say “add authentication” or “fix this error,” and it will do something. It will fix the problem. That’s great. ...

November 11, 2025 · 4 min · 654 words · Necati Demir

My 8-Year-Old Built 30 Games in 30 Days (Then This Happened)

The Setup About a month ago, I installed Claude Code on my son’s computer. By the way, I gave him a Linux computer. To be honest, I opened the terminal, installed Claude Code, and now he’s using it. I thought, “Okay, it will be cool. He’ll make some simple games; some HTML, JavaScript, and that’s it.” The 30 Games Surprise In the last 30-40 days, he coded more than 30 web games with HTML and JavaScript. ...

November 6, 2025 · 3 min · 496 words · Necati Demir

Meta's Code World Models: Understanding Code Execution, Not Just Syntax

I want to talk about an exciting research paper that has been on my list since its release last month. I finally had the opportunity to dive deep into it, and I believe this represents a fundamental shift in how AI understands code. What Are Code World Models? Let’s start with the basics and some impressive numbers. Meta’s Code World Model is an open weights large language model with 32 billion parameters. It features a dense, decoder-only architecture with a 131K token context size. While these specifications are noteworthy, the truly exciting aspect lies in what this model actually does. ...

October 29, 2025 · 5 min · 946 words · Necati Demir