Blog AI

Simple Prompts I Give LLMs to Check Their Competency

Benchmarks say the frontier models are the best in the world. Four stubborn little tasks of mine complicate the story.

I have four relatively simple tasks that I keep coming back to whenever I want to check what a language model can actually do. Currently all four fail to meet my expectations. I've tested the latest models, including Opus 5.5, which ranks at the top of several leaderboards, and GPT-6 Astra, but neither has delivered the results I expect.

The Four Prompts

  1. Find the static panel art for a specific episode of an anime. Surprisingly, Gemini 3.8 Flash via regular Gemini chat is the only one that succeeds in giving a usable response, though even it fails to find the right timestamp.
  2. Chess, Elo 1200, mid-game. Given the current board position (image attached), generate a tree of the best moves for my side, accounting for all possible opponent responses.
  3. Build me a daydream. I'm feeling detached from this world. Generate an original fantasy novel world with a well established system, rich lore, and immersive worldbuilding that I can daydream about. Remember, I've already thought of over 100 possibilities, so it needs to be something genuinely fresh and unconventional.
  4. Solve this latest JEE Advanced physics problem in a mathematical sequence. Not just the answer, but the solution developed through one specific theory and methodology, with steps suitable for a research paper proposing a new method.

How the Models Did

1. The anime panel. Almost every model gives timestamps based on international streaming metadata, which is often wrong. They then try to search the web or online communities, where they again hallucinate when identifying the correct episode and timestamp for a particular video version, especially when multiple video files exist for the same stream. As a result, none of them can provide the correct timestamp, wasting 30 to 40 minutes of my time. However, Gemini 3.6 to 3.8 Flash models perform surprisingly well. They directly tell me that a panel appears during the scene transition from scene X to Y, which is super helpful because I can immediately locate it based on my memory of the episode.

2. The chess tree. Opus 5.5 and GPT-6 Astra, in high thinking and effort mode, can identify the first few best moves perfectly. The next 3 to 4 moves are often ideal, book moves, or good moves. However, no model can generate a completely accurate tree of the best moves. Weaker models frequently suggest poor moves, mistakes, or outright blunders. Stockfish is still the way to go.

3. The fantasy world. Gemini models are good at adding creative aspects, especially when drawing on my past chat history and personal memory. However, that is where their strengths largely end. GPT models are good at providing detailed responses, but their outputs tend to be generic. Claude can be thorough, with some creativity, but its responses often feel cringe-worthy and lack intellectual stimulation. Chinese models and open weight models produce unusual outputs that are often not usable in the given context. Overall, no model consistently generates usable creative output. However, exploring the diverse outputs of 5 to 6 models might help us, as we are the ones doing the actual imagining, finding sparks of inspiration and thinking of new ideas.

4. The physics problem. Frontier models can solve most questions correctly. However, when asked to use a specific theory or methodology, or to provide steps for a research paper proposing a new method to solve a problem, their performance is terrible, with frequent hallucinations. They can be useful for validating correctness, but they struggle with meaningful cross domain connections and developing genuinely novel methodologies.

What Each Prompt Actually Tests

Taken together, the four prompts probe a fairly complete slice of model competency: tool calling and reasoning (the chess tree), web search and curation (the anime panel), writing quality and creativity (the fantasy world), and deterministic logical sequencing plus methodological discipline (the physics problem).

The pattern in the failures is consistent. Models do well at retrieving and recombining what they have seen, and at validating correctness after the fact. They break down at precise grounding (which exact version, which exact timestamp), at exhaustive deterministic search (the full move tree), at sustained originality under constraints (a world unlike the hundred I already imagined), and at inventing genuinely novel methods rather than replaying known ones.

The Takeaway

None of this means the models are useless. Stockfish still wins at chess, Gemini still finds my panels, and a spread of five or six models is still a good spark machine for ideas. But leaderboards measure what is easy to measure. My four little prompts measure whether the model can do the thing I actually need, and on that test, even the best models of 2026 still fall short.


... reads