Search jobs now Find the right job type for you Create a job alert Explore how we help job seekers Contract talent Full-Time talent Learn how we work with you Executive search Finance and Accounting Technology Marketing and Creative Legal Administrative and Customer Support Technology Risk, Audit and Compliance Finance and Accounting Digital, Marketing and Customer Experience Legal Operations Human Resources 2026 Salary Guide Demand for Skilled Talent Report Job Market Outlook Press Room Tech insights Labor market overview AI in recruiting Navigating the AI era Staffing for small businesses Cost of a bad hire Browse jobs Find your next hire Our locations

Better AI business systems start with continuous evaluation

Thought Leadership AI Workplace Research Research and insights Article
As AI becomes more embedded in workflows, leaders are under growing pressure to show measurable business value. While many may seek an AI model that serves as the gold standard, the path to better results may start with better prompts, sharper evaluation and more focused human oversight. EvalLoop: A Methodology for Evaluation-Driven Iterative Improvement of Business AI Systems, a research paper co-authored by Danti Chen, Robert Half’s Senior Vice President of ATI and Head of Data Science, draws on rigorous testing across 10 leading generative AI models to seek effective solutions. Danti sat down with us to discuss what she and her research partners learned—and why the answer surprised them.
Key takeaways  
  • There is no universal "best" AI model. No single model outperformed on every quality dimension. The right choice depends on which specific capabilities matter most for your use case.
  • Refining how you instruct AI can matter more than which AI model you use. Organizations may be able to achieve better outcomes by refining prompts before investing in larger or more expensive models.
  • Human expertise is an essential multiplier. Even the most advanced AI systems benefit from human-in-the-loop (HITL) oversight to evaluate outputs and drive continuous improvement, ensuring human insight shapes AI quality at every stage.
  • Responsible AI requires measurable, repeatable processes. Consistent evaluation and governance help organizations make informed decisions and build trust in AI-driven outcomes.
Many organizations are looking for practical ways to improve AI outcomes and demonstrate business value. What challenge were you trying to solve with EvalLoop, and why does this research matter now? Danti Chen: One of the biggest challenges organizations face today is the rapid pace of change in AI. New models are constantly entering the market, and many can seem to be the best choice for a business use case. The difficult part is determining which one best fits your specific needs. When we tested 10 models, we found that none dominated across all quality dimensions. For example, a model that ranked second overall outperformed the top-ranked model on factual accuracy by 10 percentage points. Aggregate scores hide those differences. EvalLoop helps organizations evaluate and improve prompts before comparing models. By measuring prompt quality, testing refinements and tracking changes over time, organizations can better identify whether performance issues stem from the model itself or how it’s being used. That matters because organizations sometimes assume they need a larger or more expensive model to improve performance. In many cases, refining the prompt can improve outcomes, reduce token usage and computing costs, and help organizations get more value from their AI investments.
What is prompt optimization? Danti Chen: Prompt optimization is the process of refining the instructions given to an AI model to improve the accuracy, consistency and usefulness of its outputs. Model performance depends heavily on how it’s prompted. While prompt engineering best practices exist, there isn’t a single prompt that consistently delivers the best outcomes across every model. That makes it difficult to know whether weaker performance stems from the model itself or from how it’s being used. In our research, we traced 69% of AI errors to a single diagnosable cause: The prompt was encouraging the system to make inferences beyond its source material. Once we identified that root cause, the fix was straightforward and improvement was immediate.
Why isn’t choosing a different AI model always the best place to start? Why should organizations optimize prompts before considering a different model? Danti Chen: One of the key findings from our research is that there isn’t a single best AI model. The right model depends on the business problem you’re trying to solve, and multiple models often can be effective for the same use case. If organizations focus only on finding the most capable model, they may end up paying for capabilities they don't need. A more expensive model can increase costs and consume more computing resources, but it won't necessarily improve results if the real issue is a poorly optimized prompt or another part of the AI workflow. To put a number on it: We tested a configuration change on a more advanced model, expecting improvement. But the result was not statistically significant in any quality metric. By contrast, refining the prompt based on a structured diagnosis improved the same model's performance by 15%.
What should organizations evaluate before deciding to switch AI models? Danti Chen: AI is often just part of a more complex system, organizations need a reliable way to understand what’s working, what’s not and why. A strong starting point is to establish a disciplined evaluation process. In practice, that means 3 steps: First, measure quality across distinct business areas rather than relying on a single overall score. Second, diagnose why the system is failing in its weakest area, not just that it's failing. Third, make one targeted change, re-measure and confirm the improvement before moving on. Step 2 is often where the biggest opportunity lies, and it’s the one most easily overlooked. Without a repeatable evaluation process, it’s difficult to identify what’s driving performance. A poorly designed prompt can cause even a strong model to underperform, leading organizations to replace a model when the real opportunity lies in improving how it’s being used.
What role does human expertise play in improving AI performance? Danti Chen: Human expertise is one of the most important factors in building effective AI systems. I like to think about it this way. Give a hammer to a skilled carpenter and then give that same hammer to someone who isn’t a carpenter. The tool is exactly the same, but the results will be very different. The expertise is what makes the difference. That's exactly what we saw in our research. The same AI model produced dramatically different results depending on how the prompting was refined. The same idea applies to AI. Strong results don’t come from simply having access to a powerful model. They come from understanding how to evaluate performance, identify what’s limiting results and make the right adjustments. Human expertise helps organizations determine what is actually causing performance issues. That might mean identifying weaknesses in a prompt, selecting more appropriate evaluation criteria, recognizing when a model is underperforming for a specific task or deciding when a different model is warranted. Those judgments require business context and experience that AI alone cannot provide. It's an iterative process: Change one variable, measure the impact, learn, repeat. In our case study, 3 iterations were enough to increase output quality by 15%, with the single largest gain coming from a prompt refinement that cost nothing to implement. Each of those improvements was driven by human decisions: identifying the root cause, choosing the intervention and validating the result.
What are the risks of not taking a disciplined approach to evaluating AI systems? Danti Chen: One of the biggest risks is inconsistency. Without a repeatable evaluation framework, teams may assess models differently. Over time, that makes it much harder to compare results, understand what’s working and make informed decisions across the organization. As AI adoption grows, maintaining consistency becomes increasingly important. Without a measurable improvement cycle, organizations risk creating fragmented AI programs that become harder to manage, govern and optimize over time.
How can EvalLoop help organizations improve AI-driven decision-making? Danti Chen: Responsible AI begins with having a process that’s measurable, transparent and repeatable. Organizations need to understand not only which model they’re using, but also why they selected it, how its performance was evaluated and what evidence supports those decisions. That’s why a systematic approach matters. Organizations can make decisions based on data, document how those decisions were made and measure the impact of changes over time. Technology is only part of the equation. People play an essential role in reviewing results, validating outputs and making sure AI is being used appropriately for the business problem at hand. That combination of systematic evaluation and informed human oversight helps build trust. Leaders can have greater confidence that AI systems are producing reliable results and that those results are supported by evidence.
How is Robert Half applying these principles to inform business decisions? Danti Chen: It’s important to recognize that not every business problem should be solved with AI. At Robert Half, we apply a disciplined, evidence-based approach to evaluating technology and determining what may be best suited to specific business problems. That means looking carefully at consistency, explainability and measurable impact before deciding how a technology should be used. The principles behind EvalLoop reinforce that approach. Continuous evaluation helps us understand what’s working, measure the impact of changes and make more informed decisions over time rather than relying on assumptions. Those same principles guide how we evaluate and develop AI-enabled capabilities that support clients, candidates and internal teams. Human-in-the-loop is a critical aspect of our approach, with the people closest to the problem shaping the outcomes. Our methodology is designed so that humans drive the most consequential decisions while structured evaluation handles the measurement in between. Our goal is to match the right approach to the right problem, and prove it works with evidence, not assumptions. That means AI-enabled capabilities that have been measured against relevant business outcomes and are continually improved.
Contact us As you navigate the AI era, our experts at Robert Half are here to help.