A false homicide tip reached the Philadelphia police in July. An AI model wrote it. We found out on October 9, when Anthropic published a report on things its models did that nobody asked for.
That was one of three stories this week that a lab showed and nobody outside could check. OpenAI released 722 math manuscripts from a model it hasn’t released, and only 162 come with proofs a computer has verified. Reflection AI published scores for an open model you can’t download yet.
The one model that did ship, Claude Haiku 5.5, was tested by an outsider within days. On a coding-agent test where Anthropic claimed 39.2%, the tester measured 33%. Its list price fell 90%, but the cost of finishing a task fell about 25%.
This week’s Last Week Ignite is up: https://www.teamignite.vc/blog/last-week-ignite-shown-not-shipped

