Nils Durner's Blog Ahas, Breadcrumbs, Coding Epiphanies

[UPDATED] Clash Eval

description: “Details the Clash evaluation framework for LLM comparison, describing methodology, scoring metrics, and case study results across multiple models.” layout: post title: “ClashEval: When LLM Safeguards Clash with RAG” date: 2024-05-20 last_updated: 2024-05-20 tags: [llm, rag, misinformation, ai-safety, aleph alpha] — Read more...

DeutschlandGPT: Not Quite 'Made in Germany'

A recent LinkedIn post about “DeutschlandGPT” caught my eye, promising an AI solution “Made in Germany”. Upon closer inspection, some concerning details emerged. Read more...

The Rise of AI-Generated Videos

Peter Gostev, Head of AI at Moonpig, recently shared his experience with AI-generated videos on LinkedIn. It’s a topic I’ve been following closely, and I’d like to share my thoughts on this evolving technology. Read more...

AI-Generated UI/UX: Promise and Pitfalls

From Ethan Mollick’s LinkedIn post about AI-generated sound effects to the recent Heise iX article on AI in UI design, it’s clear that AI is making inroads into every aspect of digital creation. Read more...

GPT2 Chatbot

People on Social Media are excited about a new model on the LMSYS Chatbot Arena: gpt2-chatbot. Some theorize that it may be GPT-2. I have fingerprinted the Tokenizer, and no: not GPT-2, but consistent with OpenAI cl100k (used for GPT 3.5 onwards). Peculiar gaps in world knowledge (both niche and common knowledge) were the same as in the other GP... Read more...