3 comments

  • simonw 54 minutes ago
    DeepSeek/deepseek-v4-pro and z-ai/glm-4.7 are the two models that hallucinated answers - notably, these are text input only models. It looks like the PDF was being treated as a sequence of images.

    Testing this via OpenRouter seems risky to me. It would be interesting to see results for this without a proxy in the middle that potentially confuses the results.

  • bbg2401 52 minutes ago
    > Ask two AIs and you'll get two answers. Ask the same one twice and they won't match either. The Judge gathers the best ones, makes them deliberate and rules — by the same yardstick every time, with its confidence level and a warning when even they can't agree.

    Try to speak with your peers instead of relying solely upon an LLM for product ideas. It’s obvious this product has had no human input, let alone from a subject matter expert.

  • sReinwald 45 minutes ago
    [dead]