When Kafka is the source of truth; schemas become your source of headaches

Ricardo Ferreira

Slides

Abstract

One of the coolest things about streaming systems such as Apache Kafka is their ability to handle any type of data. You can store events at Kafka and have different systems processing their event data. You may start with a few systems and add new systems as needed. While certainly possible and attractive, this isn’t simple. Schemas play a key role in how each system consumes the event data and processes them. Reason Schema Registry exists, right? Not really. Schema Registry doesn’t solve any of your data problems. It’s just a registry for your schemas. Admittedly, without it, there would be no policy enforcement. However, data problems can still happen. Issues with encoding, format mismatch between different programming languages, new code not being able to read data written by old code, etc. In this session, we will get into the weeds of data serialization with schemas. We will discuss the differences between formats like JSON, Avro, Thrift, and Protocol Buffers, and how your code must use each one of them to serialize data. It will also clarify the impact of switching Schema Registry with other registries, and whether you can use them together. If you ever wondered why your Python code can’t read something written by Java, why integers are getting confused with strings, or simply how schemas end up in Schema Registry, this session is for you.

© 2018 - 2026 Ricardo Ferreira

Search is powered by Pagefind. Just hit CTRL+K or CMD+K to start searching.

Powered by Hugo with Dream and Devrel themes.

Open Source

I contribute to LangChain4j, an idiomatic open source Java library for building LLM-powered applications on the JVM. Recent work includes adding native vector search embedding stores so developers can build RAG, recommendation engines, and AI memory systems.

I also ported RedisVL to Go, an open source, AI-native client that brings vector search, semantic caching, LLM memory, semantic routing, rerankers, and an MCP server to the Redis ecosystem for Golang developers.

Public Speaking

I’ve been speaking at conferences since 2008 and doing it full time as part of my work with DevRel since 2018. My talks go deep on the systems I build with: distributed systems and event streaming, AI engineering and vector search, and the data infrastructure that has to hold up when the demo ends and production begins. Some of the events I’ve spoken at include AWS re:Invent, Microsoft Ignite, Google Cloud Next, KubeCon, Oracle OpenWorld, QCon, Strange Loop, Kafka Summit, Pulsar Summit, JavaOne, DevNexus, JFokus, JNation, and All Things Open.

Ricardo Ferreira presenting on the main stage at AI DevWorld
On the main stage at AI DevWorld

You can find my upcoming and past talks on my speaking calendar. Recordings also live on my YouTube channel, and the code I write for talks, demos, and workshops is on my GitHub.

Consulting and Professional Services

If you’d like to hire me as a consultant for your projects, speak at your event, or lead a hands-on workshop for your team, contact me at riferrei@riferrei.com. I can understand the scope of your request and provide a free estimate.

Who am I?

I work at the intersection of AI, data infrastructure, and distributed systems, turning complex technology into things developers can understand and products users love.

Lately, that means hands-on AI engineering: building vector search, semantic caching, agent memory, and RAG into the data layer, and figuring out how to make AI agents secure enough to ship. I contribute to open-source projects like LangChain4j and RedisVL for Golang.

The AI-native work isn’t a pivot. It draws on the same systems-design foundation I’ve built for 20+ years: designing data systems for scale, moving data fast, watching where systems break; now applied to vectors and agents. I have worked on RDBMS and Big Data at Oracle; event streaming with Apache Kafka and Apache Flink at Confluent; observability at Elastic; AI and developer tooling at AWS; and NoSQL and vector stores at Redis. That foundation is exactly what separates AI demos that work on stage from AI systems that survive production.

Social Links