Why AI Still Can't Write a Good CISSP Question
Most people studying for the CISSP in 2026 have at least tried using AI to help them learn. It makes sense because the technology is everywhere and it feels like a cheat code. You request a detailed study plan and a dozen CISSP-like practice questions, and within seconds you’re off and running. Why not? The exam has a reputation for being difficult, and it can be hard to evaluate any prep platform without spending time with it.
I want to make the case that while AI can be useful in some areas for exam preparation, relying on chat tools to generate adequate exam-level practice questions may increase your risk of problems on exam day. These risks may not be obvious. AI is good at producing questions based on technical criteria but misses the mark on the type of questions you’ll likely see on the exam.
The CISSP isn’t really testing what you know
A CISSP question doesn’t normally use simple recall or fact-only formats. You’re often provided a short scenario and four options that all seem reasonable, asking which is the BEST answer or which one you’d do FIRST. People call the exam “a mile wide and an inch deep”, but the deeper truth is that its goal is to test judgment. Weaving together sometimes disparate technical, safety, cost, and business goals, a scenario can force you to make choices in the face of ambiguity. CISOs face the same sort of judgment daily.
That’s part of what makes the CISSP difficult, and what a good practice question needs to reproduce. A question that only checks whether you memorized a definition helps you train just some of the required muscle. The actual exam gives you two or more plausible answers, and your job is to pick the one that best fits the criteria. If you don’t use practice tests that reinforce that process, you’re not preparing well.
And this is exactly the part that trips AI models up.
Where AI breaks down
When you ask a model to write CISSP questions, it can fail in a few very specific and sometimes subtle ways.
AI tends to reuse words from the stem (the core part of a test question that sets up the problem) in the correct answer. In other words, it leaves fingerprints. It also makes the right answer more complete: usually a bit longer, and more qualified because it’s trying to make the correct answer airtight. These fingerprints can be subtle. Many candidates may learn to pick the most complete, familiar-sounding option and get the answer right more often than they should. That trains pattern-matching, a habit that doesn’t work well on the real exam.
The most obvious version of the AI fingerprint, where the right answer is simply the longest, is the easiest to catch. A harder one is the answer that is the same length as other options, but still gives itself away by adding more complete conditions. Of course, a word count misses this. Finding this fingerprint means examining the option to see if it is doing more work than others.
As we also know, AI makes up shit, usually with complete confidence. Ask enough questions, and the chat model will eventually invent a framework, miscategorize a control, or assign a requirement to the wrong standard. If you don’t catch it, you end up memorizing the wrong thing and carrying that error into the exam.
One of the most important issues is that AI can’t reliably understand or judge the more complex business-integrated, multi-component (or multidimensional) scenarios. It even trips on the terms FIRST versus BEST. It’s great at “which encryption algorithm fits this requirement.” It struggles with “the company just discovered x, what do you do first.” Especially where three or four options are defensible, and the answer depends on business alignment AND the right order of operations. The model will revert to the technical fix because that’s how it was trained.
This last issue is the one I’ve spent the most time on.
Teaching a machine what “CISSP-worthy” means
While building the Academy, I’ve developed an AI rubric, or “AI judge,” to rate questions like a human examiner. I write a question or a model drafts one, and the judge scores it. Is the right answer truly the single best one, or is there a defensible tie? Does it rely on judgment or a technical detail? Does it carry any fingerprints or tells?
The judge isn’t all-knowing, and it often doesn’t get it right. Current “frontier models” (Opus, Gemini, GPT-5x) all make very similar grading mistakes. They rate more technically oriented answers as harder, so a question loaded with technical distractors looks difficult to them. Even when, to a human thinking like a CISO, the answer is obvious. This is basically “hard by technicality, easy by judgment.” If AI wrote the question and then also tried grading it, it’s usually blind to its own fingerprints in the answer.
So the judge never has the final say. I use it as a detector, not a decision-maker. It surfaces suspects (“this one smells like a tell,” “these two options might both be right”), and it’s a human’s job to make a judgment call. I also pay attention when the judge expresses uncertainty. I ask it to run multiple iterations, and when its answers split, that split is something to evaluate. A question the judge can’t decide on is usually right on the border. And again, a person needs to review and make a decision.
As an example, I recently reviewed a question about securing a software supply chain, with plenty of “technical machinery.” The kind of question that feels hard. The keyed “best” answer was reasonably based on a component that was authentic and unaltered. But from a security manager’s perspective, the real question wasn’t “is the component authentic.” It’s “do we trust the source at all?” The judge rated it fine. Only reading it as a manager, not an engineer, caught that the better answer.
There’s an underlying reason and a term for this. AI is fluent at the lower rungs of Bloom’s Taxonomy: remembering and understanding. But the CISSP lives in the higher ones: applying, analyzing, evaluating.
Once you reach a certain level of technical competence, the technical distractors stop being tempting, and the correct question becomes obvious. It still may look like a CISSP question, but it doesn’t really behave like one.
None of this came together in a single clean formula. It took many versions of the rubric, and I still tune it by hand because “what makes a question CISSP-prep-worthy” is a craft, and one you can only partly automate. The judge narrows it down. But a person makes the call, and that division of labor is important to how the Academy operates.
Why I built the Academy
The trouble I kept running into, with both linear question banks and the AI-generated kind, is what led me to build a CISSP study platform called the Academy.
That “judge narrows it down, but the person makes the call” split isn’t a slogan. I use AI as a tool, but the real decisions, such as “Is this the single best answer? Is this judgment or trivia? Does this train the muscle the exam tests?” rest with a human who has taken the exam. That’s the line AI models can’t cross yet.
I looked for a single product that did all that but couldn’t find one, so I started building it. Every question is crafted to be worthy of the test bank and to build your readiness as the end goal.
Every question comes with a full explanation, not only why the key is right but why each distractor is wrong, since understanding why the tempting answers are incorrect helps build judgment.
None of it’s magic. It’s work done by hand where it must be and helped by AI where it’s safe.
The shortcut that isn’t there
The CISSP is hard because it tests judgment, and judgment is the one thing that still doesn’t get generated from a prompt. The tools got better. The people who write good questions became faster. But the shortcut, the one where you generate a stack of questions and grind them until you pass, still isn’t there, because question creation that is truly exam-prep worthy needs a human.
Use AI. It’s a great study partner. Just don’t hand it the one job it can’t yet do.



