Databricks is more than just a data platform, but most marketing teams are underutilizing it in their tech stack, or, worse…thinking of it as “just a tool their data or engineering teams use.”
In this webinar, we share 5 real use cases that span the retail, c-store, streaming, gaming, and healthcare industries — from defining requirements, to aligning key teams across the business, to building marketing solutions that drive performance in their marketing and CRM programs — that are ultimately impacting revenue.
You’ll walk away with:
- A deeper understanding of how Databricks functionality can impact marketing performance
- How marketers can create stronger alignment with their data teams
- Real-life use cases that offer a crawl-walk-run approach to leveraging Databricks for marketing
Transcript:
Welcome, everyone. Now that we’re recording, you have stumbled your way into the Stitch Databricks for Marketers webinar. Thank you so much for joining us today. We are gonna keep the Stitch content incredibly minimal today, but if you’ve never heard of us and your friend sent you this link for the first time, we are connecting data to marketing performance.
We’re a consultancy or professional services firm partnered exclusively with Databricks on the data side. Many of our backgrounds are in martech and professional services. I’ll be your host and co-presenter today, Tyler Williams. I’m a Databricks Solution Lead, and I’ll let Bobby Tichy introduce himself as well.
Yeah. Hey, everyone. My name is Bobby. I lead our Solutions team at Stitch. I’m excited to dive into Databricks and what that means for marketers.
Let’s do it. So what we’re gonna talk about today, if Databricks is a relatively new word for you as a marketer, we’re gonna lay some basic foundation for what the platform is and why you should care about it.
We’re gonna talk a little bit about your data team and how we would recommend partnering with them and better understanding them. Then we’re gonna get to the meat of the presentation, which is customer stories, real use cases, how marketers are actually using Databricks and the data coming out of it today, and then what’s coming next.
So that is a nod to CustomerLake, a recent announcement from Databricks on its agentic CDP offering. Along the way, we are gonna try to have a little bit of fun, so applause, laughter, and/or boos are all welcome on any levels of engagement that you choose to throw our way. So very quickly, what is Databricks?
You’ve never heard of this before. You’re a marketer who’s trying to learn more about data. You don’t understand why your data team takes so long to do what you ask them to do. What is Databricks likely doing in your organization today? Really three levels of maturity as it pertains to marketing.
So number one is ongoing data collection. As marketers, we have an incredible amount of sources that we’re collecting data from and ultimately generating data when it comes to customer engagement. All of that data is flowing into baseline Databricks utilization of the data warehouse, and it is stored there.
From there, your data engineering and data science teams are likely, or at least hopefully, cleaning, joining, and organizing that data. We’re not gonna talk in technical terminology today, but if you are familiar with the terms raw, bronze, silver, gold, that is what is happening at each layer of that clean, joined organization.
Gold being the primary layer of data that marketers are going to interact with and utilize because that is where we get identity resolution and usable data for the marketing team. Finally, once we have that data cleaned and joined, we can actually run models or do analytics and reporting dashboards off of that.
So think sophisticated next best action or product recommendations, that sort of thing, and then the reporting elements that come from that as well. So three levels of maturity here as we think about Databricks and why marketers should care about it. For the record, Databricks is not just a data platform.
There are powerful, real AI offerings within Databricks that we’ll speak about a little bit throughout today’s presentation as well. So Bobby and I have spent close to decades plural, we’ll go decade and decades, combined, with marketers across multiple different agencies and services firms.
And what we often hear, and what each of us probably said very early in our careers as marketers, some quotes here. I’m assuming a couple of these have either been muttered within your walls or I guess Zoom rooms in today’s remote culture. But a couple I do want to highlight here. “Why is it so hard to get the data I’m asking for?
We don’t understand what the data team is doing. Are they just lazy? What’s going on here? I need this for a campaign tomorrow for business-critical reasons. Why can’t I get that faster?” The other one that I want to hit on here is, “We feel like we need a CDP or aren’t getting value out of our CDP.”
We’re gonna talk a lot about that today. The legacy CDP promises have rarely been fully realized and delivered on from our perspective, and we think that can change with some of Databricks’ new offerings, and also what they’ve had in market for quite a while. Bobby, anything else you wanted to highlight here?
Yeah, I think the thing that we hear probably the most right now is, one, that there’s a discrepancy between Databricks and my marketing engagement platform, or whatever warehouse I’m using. So I build an audience or a segment, the data team builds for me, and then I put that into Braze, Iterable, Salesforce, Adobe, and whatever my send audience is doesn’t match up to the other one.
And so we’ll talk a little bit about that today and why that might be the case. And then the other one is, where do I start with AI? I think that with AI, everyone’s starting to try to figure out what do I use within my marketing platform, what do I use elsewhere. And as you’ll kind of hear layered throughout this presentation, Databricks is an incredibly powerful platform.
It can do a number of different things. And that’s why a lot of companies have Databricks in addition to their other data warehouses. So even if you’re leveraging Snowflake, Redshift, BigQuery as your data warehouse of record, so to speak, Databricks is still likely involved somewhere as the AI/ML layer and can help you automate a lot of your marketing operations activities.
So how can Databricks help marketers? I kind of alluded to this at the start when it comes to foundationally understanding what Databricks is and how marketers can use it, but that base layer data warehouse, right?
That’s where that data is sitting within Databricks, not technically stored, but where that data is sitting, where your data engineering teams and data science teams can manipulate it. But then as we think about moving that data through different processes, through different pipelines to drive the outcomes we want, this is really where the magic of Databricks comes into play, helping us drive reporting and data visualization efforts, automating workflows, and then leveraging different AI tools such as Genie Spaces to help marketers interact via natural language with the data they have presented to them within Databricks, all ultimately driving to — most of you probably have a CDP outside of Databricks today, mParticle, Segment, Treasure Data, etc.
All those are relevant for today’s conversation, but again, we’re gonna talk a little bit more about CustomerLake at the very end of this. Bobby, anything else you would add when it comes to this lovely pyramid we have here?
I think just kind of calling out something that I’m sure all of you as marketers struggle with — access to data, availability of data — and it’s waiting on the data engineers or the data team to get data out for you or visibility into that data, whether that’s attribution or just simple reporting or dashboards.
And just a reminder, hopefully this will make you all feel a little bit better, is you’re not alone, right? Every client that we work with has some kind of data access issues or wants more from their data. So the use cases we’ll talk through today, we’ll talk through that a little bit more.
But just to kind of levelset that: one, you’re not alone, and two, hopefully some of the things that we’ll talk through today will help you figure out how to get things to be a little bit more efficient.
There is one thing that Databricks cannot solve, and this is something that is hard to hear for some folks, but is true. It is our perspective that in the vast majority of instances, marketing teams and data engineering teams are likely not collaborating effectively. And sometimes it’s not either team’s fault, it’s the size of the enterprise.
It is product and analysts and project managers that sit in between marketing and data engineering. But it is our perspective that even if there are red tape and structural challenges with interacting between marketing and data engineering, the organizations who are delivering the customer experiences that will define the next five to ten years are already breaking down those walls, or the walls never existed.
So we want to be clear that you cannot solve a collaboration problem with strong technology. Your people have to work together, understand what you’re trying to do together, and make sure it’s not just a transactional relationship where, like we so often see, marketing needs a specific segment, they need a new attribute, and they put a ticket in, it sits there, and then they hear from data and IT when that is either complete or when they have questions because they can’t fulfill the ticket as is. So how do we understand our data team better?
Bobby, how do we do that?
Who are the characters on the screen from The Notebook?
Are you asking for audience engagement, or do you want me to tell you?
No, I’m asking you.
Noah and Allie. Not because I’ve seen this movie more than probably twice in my entire life, but I looked it up yesterday when I saw this meme that our marketing team put together, and it is a delightful one.
I don’t believe you. I bet you’ve watched this movie at least 30 times. But I think the message behind the message of this slide is that marketing and data sometimes have a hard time understanding each other. Similar to Noah and Allie, as a man, I will never fully understand women.
My wife Joni can concur with this. And as a marketer, I will never fully understand data or data teams. Every once in a while I’ll think, “I’m gonna go do a certification on Coursera or Udemy, and I’m gonna really understand data engineering at a level that I haven’t before.”
And then I get about halfway through and I’m like, “I think I’d rather just go work on marketing stuff.” So the thought behind this is really, like, when you ask for something from the data team, there’s a lot of layers behind it and what’s happening behind the scenes. I think a lot of times as marketers, we think, “Hey, I need this audience,” or, “I need this attribute to be able to personalize content or to build a new segment in my marketing platform.”
But really there’s all kinds of elements that sit behind that. We have to validate that the data engineering team has access to the source data that we want to pull. We have to make sure that the data that we’re trying to pull is actually cleansed and normalized in a way that would actually make sense to a marketer to push up into that marketing platform.
So long story short, your data team has a lot of things happening. Similar to what Tyler was saying about Databricks or any other data platform — it can’t solve the collaboration problem — don’t just throw requests over the fence. Be specific, have a discussion with them about what you’re trying to accomplish, and then drive it from there.
So have a use case that you’re trying to figure out or you’re trying to push forward, and bring that to the data team and ask them, “What would be the best way of accomplishing this with pieces of data, building an audience? Should I do this in the marketing platform?” And really work together with them.
On the next slide is really just kind of a high level of how your request from the data side actually gets done. If you see on the left-hand side, you kind of see the reverse of the pyramid that Tyler walked through. The data team’s goal is basically getting all of our raw data sources into our platform, so that we can ultimately use it in a way that drives the business forward, whether that’s for marketing, so we can set a personalized campaign, or for finance, so we can perform some kind of analysis.
Whatever it might be, their whole goal is bringing data together to be of service to the business. So when you reach out and say, “Did our campaign work?” for example, that’s a very heavy question, right? And there’s typically a lot behind that.
But what this means in practice is that when you design a journey, they’re building pipeline logic. When you ask for personalization, you’re asking for a model to score every customer that might live in the database or in the lakehouse. So there’s all kinds of elements that are happening behind the scenes, even though we might make an ask that seems fairly straightforward, or very basic.
And it doesn’t mean one team is right or one team is wrong. All we’re really saying here is that a lot of times we’re not speaking the same language. So a lot of times what might happen is that you might ask for something, the data team delivers it to you, it’s not what you would have expected, or you think it’s what you expect, but then it doesn’t perform the way you’re expecting it to.
And that’s really where that kind of translation breaks down and why collaboration is so important. Tyler, back to you.
So now we’re gonna get into the use cases. We are gonna get probably a little bit more technical, but also really focus on the storytelling and outcomes of each of these five different customers across five different verticals.
This would be a great time, whether you’re using the chat or saving questions for later, we’d love to hear from you. I think this first one, Bobby, we’re gonna talk about grocery stores.
So this is a really good example of a national grocery retailer that we work with, and they really had a couple of key issues. The biggest of them is that they had a very manual process for how they were integrating data and bringing it all together.
So before, taking more than four hours for every single time that they wanted to generate a CSV file and basically send out a personalized offer to their audience. So just for pulling the file itself was taking more than four hours, not to mention the creative that’s required, the configuration for the canvases and the campaigns that they had in Braze.
So just four hours on audience creation, and ultimately what would happen is those offers would then be outdated because it’s taking so long for us to get that data, and then all of the daisy chaining of actually being able to deploy that. So now that they’re leveraging Databricks, what they’re doing is they’ve connected Databricks with their marketing engagement platform, Braze, leveraging what’s called cloud data ingestion.
Basically what that means is you’re just connecting those two platforms through the API in near real time. So there’s no longer a wait, there’s no CSV uploads. Their team is now able to perform all the segmentation inside of Braze rather than waiting on their data team to surface that up from Databricks or wait for that manual CSV file that they’re posting.
And if you’re coming to this webinar at a three-hundred or four-hundred level — levels I never achieved in college, and I’m sure some people understand that reference — this probably seems like a relatively elementary use case.
But we find more often than not that many organizations, strong brands with what we think are great customer interactions, don’t even have the marketing data in Databricks, right? It’s sitting in a Google Sheet somewhere or an antiquated system that doesn’t give folks anything better than a CSV upload, which speaks to that four-to-eight-hour process that Bobby was just mentioning.
So they were classic download and upload before Databricks and Braze were fully integrated. So next, we’re gonna speak about a c-store or gas station chain. Again, another national brand here. I think we’ve probably all lived the life, whether you’ve worked in an agency or if you’ve worked in-house at brands, whether it’s the tenth of the month or the fifteenth of the month — generally an external consultant or maybe your internal analytics team would pull out the grand report of what happened last month, right?
So it’s July fifteenth, I’m now seeing all this amazing work that happened in June and the fruits of my labor. The bad side of that story is when the month didn’t go well, right? Certainly individual siloed teams know if paid media is doing well or if email marketing efforts are being hit. But generally as we roll that up through teams and into senior leadership, it’s hard to get a full picture unless someone’s consolidating, oh, my second PDF reference of the day, that very large deck, right?
So oftentimes many organizations are not getting real time, and certainly the marketers do not have the ability to self-service real time. Certainly if your business is highly transactional or high volume, you’re probably getting daily or weekly reports, but you may not be able to pull it yourself.
So that brings us to this story, which is this idea of moving data visualization and reporting into Databricks. Think most organizations today (Tableau, Looker, Power BI) you probably have access to a few dashboards. Are they reliable? Possibly questionable. How often are they refreshed? Nobody knows.
But this organization made the choice to move all reporting and visualization into Databricks, and then gave their marketers direct access to the entire reporting suite within Databricks, right? As we think about the challenges of going to ask your data team, “How did last month go? How did last week go? How did the last hour go?” Marketers were then able to self-serve.
And with the time saved, and this is the best part of the story, so as you’ve tuned my voice out, the next fifteen seconds are really cool, the best part about this is with the time saved, marketers were actually able to spend time looking at the data and adding additional sub-segments into the cohorts they were marketing to, and this group decided to target hot foods.
It was a very high margin item for them to sell, and they identified six sub-cohorts using this data that they were quickly able to target with loyalty points for the most engaged, big-time discounts for folks who had never bought a hot food item, that actually drove real business results. So as we clamor to think about how we can use automation or AI to save time, this is a real use case where a marketer had more time to think and strategize, get their hands on the data, manipulate it to a degree themselves, and actually identify these segments.
Pretty cool stuff.
The next one is all around personalized recommendations. I think the one element here is lean on your data and engineering team to help you drive recommendations, especially through Databricks, because it’s a very powerful way of being able to drive personalized recommendations, not just generic recommendations.
And this is really useful, especially for those in the media and entertainment industry. For those of you who are familiar with any kind of streaming service, or television or movie service, there is a law called VPPA. It’s the Video Privacy Protection Act. Basically what that law means is that you’re not able to take the viewing history and share it outside of your organization.
Legal teams have determined that basically means we’re not gonna send any kind of viewing history outside of our own storage, which means I can’t send it into a marketing engagement platform. So that becomes really difficult as a marketer because I can’t say, “You’ve watched this, so go watch this next,” or something along those lines.
All that has to happen at the data layer inside of that company’s storage. And quite frankly, I’m a big fan of this law. I don’t want everyone knowing that I’ve watched every single Avengers movie at least thirty-seven times. But outside of that, as a marketer, it makes life really difficult.
So before, when we were trying to drive these recommendations from the marketing platform, everything was generic, right? It was basically what was the most watched in total over the course of the last week, or what titles were coming out this week, things like that. So now we’re able to take all of the viewing history that we have inside of Databricks.
We have a predictive recommendations model that is then being surfaced up at the time of send at the marketing engagement platform layer. So now, since the platform knows that all I do is watch Arrested Development and Seinfeld every day for the last five years, they can recommend something like it or just get me to continue watching the episode that I left off on, rather than sending me recommendations that are generic for a horror movie like Obsession that I would never watch.
And I think this one is really important to think about the persistence that this team showed to get these use cases launched. As Bobby mentioned, this was an ongoing conversation with legal for years, around that VPPA law, right? This client had a very conservative approach to that. But we kept thinking about and solutioning — okay, well, what if the data goes here, doesn’t go here, zero-copy features, etc.
And finally got this to a point where we were actually able to launch a use case that legal approved. So I think my main takeaway here is that even when legal compliance, highly regulated industries — we all know where the line is, but sometimes legal’s gonna take a slightly more conservative approach, and that creative solutioning can actually get us to those more sophisticated use cases.
One of the things that Tyler mentioned earlier when he was talking about analytics and reporting is being able to leverage dashboards and visualization inside of Databricks, which is a really powerful element. But it also gives you the opportunity to reduce tech debt as a marketer, right? So instead of having Tableau, Power BI, or Looker, or all these kind of ancillary platforms, we can start to consolidate and leverage Databricks for more marketing activities.
One of those things is actually being able to talk to your data or talk to your campaigns. Inside of Databricks, their AI capabilities are, for the most part, called Genie. You can think about Genie the same way you think about Claude or ChatGPT or Gemini, in that you’re able to talk to it just like you would any other LLM in a chatbot interface.
But the really nice thing about it is you’re actually talking to your own data or your own organization. So for example, if you want to be able to ask what was the most impactful campaign based off of revenue, not just open rate or click rate, you could do that inside of Databricks. If you are struggling with, “What should my campaigns be in Q4 of this year?”
Or, “I need you to help me build a content calendar based on what we launched last year, but with these slight variations based on the products that we’re launching in Q4,” it can help you build all of that, not only based on all the historical knowledge it has across your company, but then obviously publicly available information, just like any LLM would be.
So if you think about resource-augmented generation, RAG, it does it at an unbelievable level because it’s specific to your company and your company’s entire data set. The other thing to keep in mind too when we’re talking about Genie or AI tools is that Databricks can be your agentic workflow motherboard, essentially.
So for example, a lot of folks use n8n, or you might be using Claude Code, or different elements like that to kind of do different things and pieces, but Databricks can be the centralized place for that. The really nice thing about Databricks being that centralized place is that internally, you basically have all of your LLMs configured to Databricks already, so any kind of enterprise agreement that you have in place will just read directly from that agreement.
So that way, you’re not having any kind of rogue licenses out there for individuals. You’re all doing it centralized. Things like automating testing for emails, or integrating Figma or Jira into your marketing automation platform — all of those different things can be powered inside of Databricks.
Hey, Bobby. First-time caller, long-time listener. We got a really good question today that I’m gonna re-ask and steal. I won’t even give the person credit ’cause that’s just who I am. How does Databricks Genie differ from something like Braze Operator?
Yeah, that’s a great question, especially because I think as marketers we’re trying to figure out we wanna use AI, but we’re not exactly sure where to use it. For something like Braze Operator, or if you’re using Iterable or Salesforce or Adobe, just think about the AI components within that platform are going to be very specific to that product. In a lot of situations, that can be incredibly helpful. It can help you build a campaign. It can provide data analysis based off of the data that you have in that platform.
The main difference with Databricks is that you’re gonna have a larger treasure trove of data because you don’t push all of your data into your marketing platform. And it’s also gonna be able to do things for you that are not just specific to that platform. So for example, we talked about an agentic workflow.
If I wanted to be able to build a reporting agent that every single time something happens that’s above a threshold or below a threshold, I could configure that inside of Databricks. So when you’re thinking about something that’s very marketing-specific or product-specific, that’s where those individual AI tools work really well.
When you’re trying to do something that spans multiple platforms or try to automate a process to make it more efficient, that’s really where Databricks shines.
Excellent answer, sir. Thank you. Let’s talk more about you, Bobby, and how my fake company that’s actually a real company is gonna market to you more effectively. So this comes back to the unified customer profile, right, and alludes to the long-term offering that Databricks recently launched around their agentic CDP.
But this customer had this vision long before CustomerLake was ever announced, was ever a thing, right. That truly is using Databricks as a homegrown custom CDP offering to consolidate all these different data sources. This is a large customer, large brands, many different divisions, products, etc.
And what they found was that if Bobby bought three different products from them or from three different divisions, their marketing and their customer communication, their customer service just became Bobby A, Bobby B, and Bobby C, right? Which is frankly probably the expectation for most of us nowadays if you bought three different products from a large brand, “I’m just three different transactions to you.”
But the end goal and state here is the idea that, “Hey, if these happen within a similar timeframe or over time, instead of Bobby A, Bobby B, and Bobby C, it can be Bobby ABC.” This customer consolidated eight different end source systems into Databricks to identify that golden record, all while also consolidating five different messaging platforms into a more modern customer engagement platform.
Kudos to this customer who’s doing both those at the same time. Massive project, ended up being highly successful. But what that gave them the ability to do is long-term make sure that they are consolidating messaging. So again, that Bobby ABC within the same dynamic email template perhaps, or prioritizing.
Saying, “Hey, we would’ve sent Bobby all three of these messages. We’re gonna prioritize this one out of these three.” What that does when it comes to your customer engagement platform, it has the ability to decrease your cost because you’re not sending as many messages, and you are likely also delivering a better customer experience through that prioritization.
What it also does for them, it allows them to identify cross-sell and upsell opportunities, from a product A to product B, from a maturity or usage of those products. And that is something that this customer is looking forward to taking advantage of very soon. So really exciting, both cost savings and incremental revenue increases. It’s always fun when you can do both of those at the same time. That’s what I’m told, at least.
Jim just put a question in the Q&A section. How about using a Slack bot with Genie via MCP so that questions to the lakehouse and CustomerLake can be asked by non-database users and executives who are in Slack all day?
Yeah, Jim, that doesn’t seem as much of a question as it is just a good idea, because you’re exactly right — we can absolutely do that. I think that whether you’re using Slack, Teams, whatever it might be as your internal chat tool, you can absolutely connect to Databricks so that the folks might not even know that they’re actually talking to Databricks, but they can talk to their data in a way that doesn’t require them to have a Databricks license or to actually log into a Genie space or anything like that.
So great observation, for sure.
Yeah, the only thing I would add to that, it is a great idea. I’m sure we’ve all experienced when senior leadership goes down a rabbit hole, whether they meant to or not, and they can then create a little bit of chaos. So with Genie spaces, you can build those out to reference certain parts of your data set, or obviously you can open it all the way up.
So I would be thinking through, when it comes to the remit of a leader or an individual, what that looks like from a governance perspective. But it’s certainly a great idea. Nothing further. Great question, great comment.
So we talked about five customer stories. We sat on those slides for a while. Single slide with the five things that you can take away. If your boss asks how that webinar was, or as you’re leaving a Yelp review for Stitch webinars on Yelp, here’s your slide: the applications of Databricks for marketers.
So basic data integration. Number two, self-serve analytics and reporting that will give time back and be more real time. Third, personalizing recommendations and working around sometimes compliance and regulated industries. Four, “I don’t know where to start with AI, I just use ChatGPT to build PDFs, and my marketing team gets mad at me” — a Genie pace is an awesome way to utilize both for marketers and, as was mentioned previously, open it up to the organization to interact with. And then finally, utilizing that golden record.
We’re gonna talk here in a minute about traditional CDPs versus the CustomerLake roadmap. And that is certainly gonna be a conversation that gets very interesting over the course of the next six, twelve, and eighteen months.
So here are your key takeaways. This is the one to screenshot. So Bobby, Databricks CustomerLake.
Yeah, thanks. And Renato actually put a question in the Q&A section. “If I already have a substantial number of data products developed in BigQuery and Databricks, should I focus solely on Databricks as my CDP, or is it necessary to use an additional tool like Hightouch?”
We’ll actually talk about that here in the next couple slides as we talk about CustomerLake, and we’ll answer your question specifically around whether or not you need Hightouch after we talk through these additional slides. I think the thing that most of us are dealing with as marketers is that the problem CDPs were built to solve was that customer data lives everywhere.
You’ve got website behavior, mobile, e-com, point of sale, loyalty, customer service. There’s no single person, or kind of record, of all of these different things. So ideally, what a CDP was meant to do was, one, to collect all of this data, unify it into a golden record, so that we could use it for segmentation, but then also customer service, or finance, or whoever needed access to that individual profile would have access to it, and then be able to activate it downstream, whether that’s in an ad platform or an owned channel.
And I think what we all kind of realize at this point is that after — especially for those of you working in B2C situations — you’ve probably gone through two, maybe even three CDPs at this point, that we kind of think of as traditional CDPs. So it might be Segment, Treasure Data, Tealium, mParticle, Data Cloud, etc.
The promise was that we were gonna have this golden record, everything was gonna be combined in one place. Well, as you all know, what ended up happening was that the implementation of those platforms was never a set-it-and-forget-it type of situation. Those companies did a phenomenal job selling us on this being easy to do — getting all of your data into one place — but we needed just as much, if not more, data and engineering resources to get that data from those sources and systems into the CDP, in addition to that also needing to go into the data warehouse.
And really what’s happened over the last three to five years is the data warehouse has become the place where all of that data starts to land, and that’s where we want to try to unify it and govern it, rather than it being in a separate standalone or traditional CDP. On this next slide, you’ll see kind of the two high-level architecture diagrams and mainly why CustomerLake inside of Databricks is different.
The one area where that standalone CDP is — you’ve got to build all those connectors. Even if they do have a productized integration, it’s typically not foolproof. We have additional elements from those connectors that we need to build on. Then we have all of that into our traditional CDP that’s gonna have its own proprietary data model, right?
It’s gonna have attributes, events, relational data, all those different kinds of things that we have to remodel. And then we want to activate on it, but then we also need to send that data back to our data warehouse, so that we have it for historical purposes too. Now, on the other side of that is where you have CustomerLake, which Databricks announced a couple of months ago.
CustomerLake is an agentic CDP, and basically what that means is that every feature or element of the product is based off of agentic workflows. So it is constantly learning and making better decisions the more it understands what you’re trying to accomplish. To put it simply, you don’t have to create new integrations.
You don’t have to pour or push all of that data into a separate CDP or a separate data model. All of that’s happening because CustomerLake is embedded inside of Databricks already. There are a number of different helpful elements to that, but one being data engineering has already done all this pre-work of unifying our data or getting it all into a single place. We don’t have to do that all over again.
There are three main components to CustomerLake. One is identity resolution, both deterministic and probabilistic. Two is segmentation, so that we can leverage all the data we already have inside of Databricks. And three is what are called Infinity Campaigns, and you can think about this as decisioning at a marketing campaign layer.
What this allows you to do is take any number of campaigns that you might have in your marketing engagement platform, or inside of paid media audiences or lookalike audiences. You put all of that into one Infinity campaign, and that Infinity campaign decides what content and what channel and what message I should receive based off of all of my interactions.
So obviously that includes — if you’re, for example, let’s say you’re a brick-and-mortar and an online retailer, that’s gonna take everything into account because all of that data is already sitting in Databricks. Or if you’re a streaming service, it’s gonna take into account every piece of data we have — how long you’ve been a customer, have you ever attrited, what are you watching — all those different things, because we have all of that data already in Databricks available to us, and we don’t have to reformat it. So you can start to see where it provides a lot of value as a CDP, because it’s not sitting outside of the warehouse as another bolt-on.
Now to the specific question: should I focus solely on Databricks as my CDP, or is it necessary to use an additional tool like Hightouch? It really depends. I know it’s a very consulting answer, but it depends on the use case. Now, if you’re using Hightouch for identity resolution, you no longer need that because CustomerLake will serve that for you, both deterministically and probabilistically.
If you’re using it for segmentation, you’re not gonna need it anymore because you’re gonna have the segmentation element of CustomerLake. If you’re using it to integrate into downstream systems, that might be a reason why you may want to keep it for a short period of time as CustomerLake continues to build out their catalog of integrations, but they’re starting with a large amount of those to begin with.
So really, I would look at what you are looking for the CDP to accomplish, and then work backwards from there, because we don’t want to just make generic technology decisions without understanding what it means for the business. Jim, your question — can Infinity Campaigns manage the multiple ICPs for B2B buying groups?
This is a really good distinction. Infinity campaigns are at the individual level, based off of the identity resolution that you set up. So if there are multiple ICPs for a B2B buying group, let’s say that you’re targeting a company and there are 12 key contacts, if all of those key contacts have different unique identifiers, which ideally they should, then each one of those could get a different message based off of those Infinity campaigns.
So maybe you want to send one person to a landing page to download a white paper. Maybe someone else you want to try to schedule a call with. Maybe another person you want to try to get to a webinar. All of those different elements could be put into an Infinity campaign, and each person would get a different message based on what they should receive.
I think the exciting thing about the potential to deprecate standalone or third-party CDPs — we’ve all been in those meetings where someone has to walk through, “Well, that field is named this in the source, it’s named this in the data warehouse, and then we actually had to change that and remove it, so it actually says this in the CDP, but it lands in Braze on this attribute.”
We’ve all, I’m sure, heard those meetings where there’s one person in the basement of the office building with a red stapler, and they’re the only one who actually knows why data fields are named the way they are. This should help to consolidate and simplify some of that. So where do we go from here?
This is our last slide here. So how we’d recommend getting started today when it comes to a couple of things we’ve mentioned around both collaboration and marketing access to Databricks. First thing is work to build direct relationships with your data team. It’s kind of a cliché, which means there’s some truth to it and some not.
But there is some truth to the idea that data and engineering teams have probably built some pretty cool models that marketing doesn’t even know exists. Oftentimes we’re very transactional in nature. Marketers want to lead the way with their strategy or their brand campaign ideas. Go ask your data team what they’ve built for maybe even other departments that you could leverage.
I’ll give a very short-winded example here. We were working with a sports team last week who was talking about how they were modeling game attendance out throughout the season with all these different variables — who they were playing, what time the game was, if they would still be in playoff contention.
And the question that some smart person asked was, “That’s awesome. All those models are getting to the marketing team so they can increase paid media spend, or partner with the community to do a community night, or two-for-one ticket sales or something.” And the data scientist goes, “No, we just give that to the revenue team so they can understand how bad it’s gonna get towards the end of the season.”
Oh, that’s a bummer. So work with your data team, get to know what they have. Get your marketing data into Databricks — I hope most people on this call have it; if not, important starting point. If your data is in Databricks and it is consolidated and clean in the ways we’ve talked about today, have your team build you one self-service analytics use case.
It could be campaign-based or some other use case based off of what you run in marketing, whether that be paid media, email, ongoing CRM, etc. One self-serve real-time analytics use case. And then finally, Genie spaces are relatively easy to set up. Start interacting with your data, right? I joke about ChatGPT and Claude, but my assumption is lots of marketers are using AI outside of a lot of the context within their business, right?
And the AI is only as strong as the context it’s served. So get using Genie within Databricks. As we wrap and move to Q&A, Bobby, any final thoughts on getting started otherwise?
I think the biggest thing is building relationships with your data team. First do it as a kind of workshop, and then make sure there’s an ongoing cadence.
So if you have never met with your data and engineering team in any kind of functional way, I would start with having a lunch-and-learn or something similar where you’re just asking them questions. Ask them to come with what are the models and all the things that you’ve built in the last three months, so that you can get a good understanding of what’s available already, because that’s gonna spark a lot of campaign ideas.
Not only net-new campaigns, but how you could optimize current campaigns and how you could personalize messaging more as well. And then build out an ongoing cadence of that. Maybe it’s a monthly show-and-tell where you show the data team what you’ve been working on, and they show you what you’ve been working on, and you work together to figure out how you could make that more impactful.
And then we just got a question from Ross. Hey, Ross, thanks for joining. Do you all have estimates on the costs of running Profile Agent? That’s kind of the nice thing about Databricks, but also the somewhat complicated thing about Databricks: everything in Databricks runs off of consumption.
So while there’s not a specific corollary of, like, “X” — this, like, Profile Agent or Campaign Agent roughly relates to this price — it’s basically about the number of records or how much you’re consuming to get to an endpoint. For example, if you’re only running 1,000 records through Profile Agent, it’s gonna be substantially cheaper than if you’re running 10 million through that as well.
But your Databricks team should be able to help you estimate those costs based off of the size of your database.
Yeah. Yep. Whoo, that was fun. The other thing I would add to that would be that your team certainly has different identity resolution elements in place, from raw to bronze to silver to gold, right? Profile Agent can certainly be leveraged as the pinnacle or key element of identity resolution. I do imagine as organizations start to adopt CustomerLake, and certainly Profile Agent, that maybe we start out small with specific use cases, and the idea there is that marketers are leveraging that, right?
So maybe you work in an industry where, hey, we have to have absolutely perfect resolution from a compliance perspective, and we have consistently seen that what our data team puts together only gets us 99.2% of the way there, and therefore we can’t trust who we’re actually messaging.
Profile Agent may be a really good and relatively cheap way to address smaller cohorts or testing specific campaigns. But on a broad scale, yeah, I would work with your Databricks rep on broad estimates, certainly.
And we’ve got another question, from Emily: “I don’t work closely with my data team today. How can I best approach them? Where should I start?”
Bobby, you kind of mentioned it. Free food always gets people excited. If you don’t work in an office setting, though, I think just a very casual introduction — and oftentimes just even the non-business answer, starting to build a personal relationship, right? I think a lot of times — I’m not often accused of being very similar to Simon Sinek — but oftentimes I think we miss the why when it comes to why another team is behaving a certain way.
So I think starting with that human connection and getting an understanding of, as a marketer communicating to data, “Hey, when I ask this, here’s why I’m asking for this,” or, “I know I’m never giving you enough lead time, but when I ask for a list pull on Thursday afternoon, it’s because sales have been down this week and we’re trying to do a big push for the weekend. So how can I work with you better to understand what I am being pushed by the business for?” So I would make it human, vulnerable, all that kind of stuff when it comes to building personal relationships. And then, as Bobby mentioned, ongoing sessions, whether that’s one-on-one lunch-and-learn sessions or whatnot.
Bobby, anything else you would add to that?
No, I think you nailed it.
Great. Thank you. I don’t believe we see any further questions, therefore we’re not gonna give any more answers. So anything else in the chat? I’m not even sure if people can come off mute to talk to us.
Okay, we got another one in the chat here. I will read it, and Bobby, one of us can take it. In the roadmap, are there plans to build an activation layer to trigger email, SMS, and push notifications?
At this time, not that we’re aware of today. The biggest thing that we saw during the Databricks launch of CustomerLake is that they are heavily partnering with all of the marketing engagement platforms, the main ones that you are aware of.
So Braze, Iterable, Salesforce, Adobe are kind of the main four that they’re starting with as activation partners. I don’t think that, at least from what we’ve seen on the product roadmap and what we’ve heard from that team, there is a plan to actually get into the content management or the actual message activation layer.
One follow-on I would add there, maybe Databricks ultimately does that in years. I think the pending competitive element, or battleground, will be the decisioning logic, and even to a degree orchestration, right? Because I think that’s where CustomerLake starts to get into, certainly with Infinity Campaigns, and whether you’re using Marketing Cloud, Adobe, Braze, etc. — you’re likely doing a lot of that today with some level of it being done within your CDP or Databricks or whatnot. I think Databricks CustomerLake is squarely pushing into that ongoing decisioning area that will start to be competitive with your customer engagement platforms.
You’re welcome. All right, team, with that, I’ve completed my awkward pause. We want to say thank you for your attendance, questions. Really appreciate you registering, joining. Recording will also be made available.
So thank you very much. Go work with your data teams better. If you’re a marketer, get yourself into Databricks. And if you want to learn more about Stitch, or Bobby and I, the horrific movies that we enjoy making jokes and quotes about, you should be able to find us relatively easily — LinkedIn, etc.
So thank you very much, and we’ll chat again soon.