MIT FutureTechSeptember 16, 2026
Mihai Codreanu, Alex Imas, Juan Mateos-Garcia, et al.
Why it matters. Triangulating 15 million Gemini interactions, more than 2,600 specialized models, and a survey of more than 600 scientists, the study finds researchers saving close to seven hours a week while reporting heavier verification demands and a growing backlog of hypotheses they cannot get through. The time did not leave the system, it moved. Hours returned at the production end arrive downstream as checking, prioritizing, and deciding what is worth pursuing, which is judgment work and the hardest kind to staff. Hire against the bottleneck you are about to have, not the one that just cleared.
Boston Consulting GroupSeptember 24, 2026
Sagar Goel, Julie Bedard, Matthew Kropp, Ashley Sim, and Charikleia Kaffe
Why it matters. A qualitative study across 50 companies finds the front-runners redesigning work around outcomes: AI produces, people direct and evaluate, and a team owns the result. What separates them is not model access, it is decision rights. Who approves, who answers for a wrong output, and what the workflow does with a disagreement. Those questions have organizational answers and they get settled before any tool arrives. A company that has not answered them buys the same model as everyone else and gets less out of it.
arXiv preprintSeptember 15, 2026
Robin Welsch, Michelle Rausch, Daniela Fernandes, et al.
Why it matters. In a 535-person reasoning study, assistance beat unaided work while the human and AI pairing did not reliably beat the assistant on its own, and participants followed the advice at times when it was wrong. A human in the loop is a seat, not a safeguard. What converts one into the other is calibration: knowing which claims to check, which to push back on, and which to let stand. An hour of observed work shows you that. No credential a candidate can send you does.