Prashanth Harshangi is co-founder of Enkrypt AI, building security and governance infrastructure for enterprise AI deployments.getty​Ask a leading large language model whether one caste is more fit for leadership than another and it refuses to answer. Ask the same model to fill in the blank in "the ___ caste is associated with purity" and it completes the sentence with a stereotype. Same model. Same identity groups. Two completely different bias profiles. Most enterprise bias audits would catch the first behavior and miss the second entirely.That gap is not a corner case. It is the central finding of a study my team at Enkrypt AI Research Labs recently published, auditing seven commercial and open-weight large language models across 45,000 prompts spanning nine bias types and seven evaluation task formats. The takeaway for enterprises is uncomfortable: Most bias evaluations measure one slice of a model's behavior and report it as the whole picture. Procurement decisions are being made on incomplete data.What We MeasuredMost published bias evaluations test fewer than 1,000 prompts across one or two task types. We built the audit to span the full range of how a model can be asked to reveal a bias, from explicit decision-making ("which candidate should be hired") to implicit association (fill-in-the-blank, sentence completion, analogy). We covered axes that mainstream benchmarks tend to under-test, including caste, linguistic and geographic bias. We ran the same prompts across multiple models so that any divergence reflected the model, not the question.The reason for the scale was simple. We wanted to know whether a model's bias score on one task predicted its bias score on another. The answer, with some uncomfortable consistency, was no.Three FindingsBias Is Task-DependentFor the same model and identity groups, stereotype scores diverged by up to 0.43 across task types. A model that looked well-aligned on explicit probes reproduced stereotypes on implicit ones. This is the finding that most directly undermines current evaluation practice: A single bias score is not a property of the model. It is a property of the model combined with the task used to measure it.We are not alone in observing this. Independent research from Princeton and the University of Chicago found that models that had improved on standard bias benchmarks, including GPT-4, still showed stereotype associations on implicit probes modeled on the Implicit Association Test. Passing an explicit benchmark tells you what a model refuses to say, not what it still associates.Safety Alignment Is AsymmetricModels reliably refuse to assign negative traits to marginalized groups. They also reliably assign positive traits to privileged ones. The first behavior is the visible work of alignment training. The second is what alignment training has not been pushed to address. The result is a model that looks safer than it is: The most measured failure mode, overt negative stereotyping, has been suppressed, while a quieter one, reinforcement of positive stereotypes about dominant groups, continues unchecked.Understudied Bias Axes Show The Strongest StereotypingCaste, linguistic and geographic bias produced the highest stereotype scores across nearly every model we tested. The implication is that alignment effort tracks benchmark coverage rather than harm severity. Where evaluators have built strong public benchmarks, models have improved. Where benchmarks are weak or absent, behavior remains close to the pretraining baseline.Others' findings point in the same direction. This study argues that biases tied to non-Western identities remain structurally under-evaluated even as these models cross geographical and cultural boundaries, and demonstrates empirically that they persist in a leading commercial model.Why This Matters For ProcurementFor an enterprise comparing AI vendors, three implications follow directly.First, a bias score from a single benchmark is not a bias profile. It is one data point on a multidimensional surface. Two models with identical scores on a popular fairness benchmark can have substantially different behavior in deployment, particularly if your use case involves implicit reasoning or covers identity dimensions the benchmark did not test.Second, models that score well on U.S.-centric fairness benchmarks may have the worst behavior on the axes most relevant to a non-U.S. customer base. If your product serves users in South Asia, the Middle East, sub-Saharan Africa or any region where the dominant bias axes differ from those measured by the leading public benchmarks, your vendor's headline scores tell you very little about what will happen in production.Third, single-task evaluations compress the apparent gap between well-aligned and poorly aligned models. When we measured across task types, models that looked similar on one benchmark diverged sharply on others. Choosing between vendors on aggregate fairness scores assumes those scores are stable across context. They are not.What To Ask For InsteadThree concrete asks will materially improve the quality of the biased information you receive from vendors.Ask for results across at least one explicit task type (decision-making) and one implicit task type (association or completion). The divergence between the two is more informative than either score in isolation.Ask which identity dimensions and geographies the bias evaluation covered. A model evaluated on U.S.-centric race and gender benchmarks has not been evaluated for caste, language or regional bias. If your deployment touches any of those, the vendor should be running those evaluations or you should be commissioning them independently.Ask for the task-by-task breakdown, not just the headline score. The aggregate number hides exactly the variance that matters for risk assessment.The Closing ThoughtThe model that refuses the question is not always free of bias. Sometimes it is just the model that knows when it is being asked. Distinguishing between the two requires more evaluation than the industry is currently doing. The gap between what is measured and what is shipped is the procurement risk most enterprises do not yet know they are carrying.​Forbes Technology Council is an invitation-only community for world-class CIOs, CTOs and technology executives. Do I qualify?