Beyond magic promises: Implementing productive AI
Omnichannel Podcast Episode 43
Watch now
Listen Now
In this episode, Noz Urbina interviews Ilya Venger, Data and AI Product Leader at Microsoft, to deliver a masterclass in practical AI implementation for business leaders.
Ilya addresses the trillion-dollar question facing every executive: Should we build our own AI solution, buy off-the-shelf, or wait for the technology to mature? His answer: it depends on understanding your specific business problems, not chasing shiny technology.
Key Takeaways
The 80% Solution: Ilya reveals that AI systems work correctly about 80% of the time. Success isn’t about perfecting that last 20% through expensive fine-tuning – it’s about redesigning processes to work with AI’s probabilistic nature. As Noz puts it,
“If you create a workflow with zero tolerance for error, you’ve designed a bad process.”
The Fine-Tuning Trap: Ilya shares cautionary tales of companies spending millions to fine-tune models for specific problems (like the “six finger problem” in image generation), only to watch base models solve these issues within 18 months. His stark example: a model fine-tuned to be cheaper than GPT-4 became pointless when GPT-4’s price dropped tenfold.
Data Reality Check: Both speakers agree that most organizations have “data heaps” – disconnected silos without understanding or metadata. Ilya’s metaphor:
“You’ve got gold nuggets in a dark room. You need to turn on the lights first.”
Organisations must understand their data landscape before implementing any AI solution.
The Build vs. Buy Decision Framework:
- Build (Fine-tune): Only when you have extremely specific tasks with proprietary data (like recognizing manufacturing equipment or crop diseases)
- Buy: For most use cases, using off-the-shelf models with good system prompts and workflow design
- Wait: When your problem might be solved by next quarter’s model improvements
“I think implementing AI for the sake of implementing AI is the wrong approach right from the beginning. Yes, you want to implement AI because you want the shiny toys, but really you need to understand what business value you’re trying to drive.” – Ilya Venger
“Right now we’re adding AI into existing processes, which is not great because we’re inventing horseless carriages instead of cars.” – Ilya Venger
What you’ll learn
.
Notes and instructions
Books mentioned
LinkedIn profiles mentioned
Chapters
00:00 Introduction and guest overview
01:42 Understanding AI implementation challenges
04:54 Metaphors and mental models for AI
07:32 AI failure modes and process integration
11:42 Advanced AI strategies and fine-tuning
31:31 Large vs. Small Language Models
33:08 Challenges of fine-tuning small models
36:21 The six-finger problem and AI evolution
44:45 Preparing content for AI and Knowledge Graphs
52:02 The future of AI and data management
Full session transcript
THIS IS AN AUTOMATED TRANSCRIPT
[00:00:00] Noz Urbina: Hello everybody. Welcome to this episode of the Omnichannel X podcast. Thank you for joining us again, or welcome if this is your first time. If you don’t know me, my name is Noz Urbina. I am founder of Urbina Consulting, your podcast host, and a very enthusiastic person who’s going to be talking to Ilya Venger today, principal product lead for industry AI at Microsoft.
Here at the Omnichannel X podcast, we explore how organizations can build better relationships with their audiences using content design, governance and systems. So of course, we talk a lot of strategy and emerging technologies, which means plenty of AI.
Ilya is a PhD in systems biology and bioinformatics from the Weizmann Institute, and he was previously executive director at UBS, where he built the enterprise Knowledge Graph Helix. He currently leads global teams delivering generative AI platform solutions across Microsoft industry clouds and is focusing today on best practices and implementations for businesses implementing AI, so he’s a perfect person to be on Omnichannel X. So Ilya, anything else you want to add about yourself?
[00:01:14] Ilya Venger: Very excited to be here. Thanks for inviting me. I’ve been the last four years at Microsoft. Very interesting times, as we all know, working first in classical machine learning and now incorporating Gen AI for the last two, three years. Very interesting journey with a lot of different customers from different industries, starting from retail, manufacturing, healthcare, financial services. As you said, we’re Microsoft industry cloud, so there’s multiple different industries and industry-specific solutions that we are delivering in those areas.
[00:01:44] Noz Urbina: I know that a lot of people in the OmniX audience come from different industries, a lot of business to business as well, although not exclusively – there’s definitely major retailers and so on, but mainly enterprise. So I think it’s going to be quite apt for our audience.
So Ilya, to kick us off: Many businesses feel pressure to implement AI, but don’t know where to start. This is one of my favorite questions. Where should AI fit in existing workflows? There’s a lot of talk about people replacement and so on, but I think it’s good to talk about whatever we’re doing, we’re going to be placing it into human workflows as well. So how do we make this integration work?
[00:02:27] Ilya Venger: I think implementing AI for the sake of implementing AI is the wrong approach right from the beginning. Yes, you want to implement AI because you want the shiny toys, but really you need to understand what business value you’re trying to drive with your business and where it could help you.
One of the most important things that we do, as we are getting into all of our customer engagements, is starting to understand: What are the main pain points? Where can you see some automation or some augmentation of your people? Where can it be helpful? And then looking at the specific processes in which people participate.
The models have been improving significantly, and the tooling around the models has been improving over the last few years. It started with very simple chatbots, and we started adding grounding in your documents and memory. Now we’re talking about agentic flows that are self-correcting at times.
I think one of the ways to look at your problems today is saying: Where can I add – where in my process can I add an eager but relatively junior person? And the best thing to think about is, we don’t want to necessarily replace the junior people. It’s actually, where could you – if I had a little army of juniors that have something to do within my process, where would I put them such that I achieved my business results?
So you need to look at your processes. This is where you start – identify the problem, identify the process, which process that you have, and then which points in those processes will augment the process and improve it significantly.
[00:04:12] Noz Urbina: I’m smiling myself silly here, because for two years I’ve been using exactly the same metaphor – I called it the “10,000 interns.” If someone gave you the army of 10,000 interns, how would you put them to work?
Some people get offended because they feel like I’m speaking badly of interns. But I think this talks to the quality problem.
[00:04:43] Ilya Venger: It’s a real problem, I think.
[00:04:46] Noz Urbina: Good, we need to be able to talk about it openly and honestly. The way I like to talk about it is that if you create a knowledge worker flow in your organization that has zero tolerance for error, that’s a bad flow. You’ve designed a bad process. If no one can make any mistakes, you have designed a bad process – that’s not on the entities involved.
So if they’re artificial or human and they make mistakes, your flow has to be able to cater for that. Although we are looking for automation opportunities, we’re looking for automation opportunities that still have some robustness. People throw out different percentages, like the 80% that almost works, 95%, etc. How should we approach this gap?
[00:05:38] Ilya Venger: First of all, build out the 80%, and not just as a POC. Because POCs usually work at 100% just on a very simple problem – the happy path that is not taking you all the way.
Build your 80%, identify what the 20% is, and then start thinking. You need to perform some analysis. One option is to start pushing your AI to be better in the particular task. It might work. It might bring you to 95 or sometimes to 99% quality, which might be sufficient for your use case.
The important thing is to identify which part of the process fails. Because alternatively, what you should be doing is saying – and this is where we eventually are going to go – right now we’re adding AI into existing processes, which is not great because we’re inventing horseless carriages instead of cars.
What you really want is to eventually re-engineer your process, and not just one process. You want to re-engineer your whole organizational, interwoven process matrix. One opportunity to look at is: Here’s the 80% that works. How should I be changing my process in order to accommodate for that slack? Maybe the 20% you actually could completely get rid of.
[00:07:20] Noz Urbina: Let’s stop and define the problem a little bit here. When we say 80%-20%, so 20% of the time it doesn’t work – it fails in some way. Let’s talk about different failure modes. Why does it go wrong? How can it go wrong?
What I’ve noticed when giving AI trainings is that if you’re new to the technology, the failure modes seem alien, because we’re used to working with computers. We say “look this up” or “tell me this thing” and it either does or doesn’t – it bugs or returns no answer. Whereas these are machines that can return wrong answers or non-existent answers. What are the failure modes that could be in the 20%?
[00:08:26] Ilya Venger: Let’s cover two things. When I said put in your machine and see what your baseline is – what’s the 80% – you need to start with evaluations. Very often when you first come to a customer with a product, you throw together some POC. The POC probably works 100% on something very small, and you actually haven’t put evaluations in place.
One of the most important things is to understand what failure modes you have. You need to have some baseline and some evaluation of what good looks like, and what is the actual delta between what is expected and what you’re actually getting out of the box.
The mode of failure – we expect our machine to answer perfectly because we’re used to deterministic programs that define the outcome very robustly. So this is one mode of failure when the machine underperforms because it just makes a standard error.
But we need to remember that the errors it makes – we tend to anthropomorphize our AI assistant. The modes of failures are not the same as an intern. They are inhuman, they are alien from that perspective. You need to identify the type and say: Is this something that a human would do, or is it not a mistake that a human would do?
For example, one common mode of failure with RAG (Retrieval Augmented Generation) when we want to get documents and have our agent answer based on these documents is semantic confusion. It takes something from one document, something from another document, and connects it in a way that no human would. It’s not necessarily a hallucination – people hallucinate all the time.
[00:10:38] Noz Urbina: People make up stuff continuously. If we’re tired or emotional, lots of us will remember things differently than someone else. There are two things: we anthropomorphize the machines, and we also idealize the humans.
There’s an overlap there – we pretend… I had literally this dialogue on LinkedIn where someone said if a lawyer gave you two different answers, you would fire them. And I said there’s no lawyer on the planet who could literally give you the same answer twice. When a human gives you an answer, they’re not doing a lookup in a database.
[00:11:32] Ilya Venger: Let me give you another example. If I go to lawyer A and they give me one answer, I’ll go to lawyer B and they give me another answer. That doesn’t mean I need to fire all lawyers. Lawyers sometimes disagree between each other, or doctors that have opinion A and opinion B. You actually go and get a second opinion with a medical professional – that is very common.
[00:11:53] Noz Urbina: Opinion or interpretation of the same facts. You mentioned the word deterministic. We’re used to programs that are deterministic, where some human determined how they should work and gave them particular rules. When they cannot determine something, they just bug – they cannot return an answer.
With probabilistic machines, which is what we’re dealing with in generative AI, they’re working on statistically probable answers. Based on experience and information, this is probably the answer. There is a level of interpretation there.
That’s analogous but not exactly the same to how humans work. We have to admit that with us and with machines, there’s an amount of interpretation and possibility of error in recall. How do we build a process that caters for that?
[00:13:16] Ilya Venger: I’ll touch on one important point. Organizations are composed mostly of humans and some technology, with processes running on top. When humans work with humans, we’ve got theory of mind. We have some expectation of what our counterpart knows or what biases they might have because of their background or education.
So we correct based on our expectation of the answer. If I tell you something and you respond, I think about why you responded that way. Then I can correct my assumption based on that.
Most processes in real organizations are not just step by step. They’re not that deterministic. Even when we write them down as Step A, Step B, Step C, if you look at the organization, that almost never happens like that.
[00:14:27] Noz Urbina: Even within step one, within step two, there’s lots of human agency and decision making. The granularity of computer steps is extremely refined and exact – “take this string of letters and put them over here.” That’s not how you write human processes.
[00:14:44] Ilya Venger: With standard programs, we expect them not to fail. We give them very strict guardrails – you are responsible for copying from Excel cell A to cell B.
But what do we do with those failures where we understand there is a problem? I might be slightly controversial, but I would say you need to understand what the probabilistic AI LLM-based machine actually does in this spot. We don’t have theory of mind of the machine.
[00:15:32] Noz Urbina: Sorry to cut you off, I just want to define the term theory of mind. For those unfamiliar, theory of mind is your understanding that there’s another mind, and that other mind is not seeing necessarily the same facts and world that you are. They’re seeing it from their perspective.
Every time you’re in a conversation and think “maybe I’m not being clear” when someone responds in a way that doesn’t make sense to you, and you reflect “Oh, maybe I didn’t phrase my question properly” – that’s theory of mind.
We need to develop, if we’re going to be using AI, a theory of how can these AI fail? Because it’s not quite like us, but it’s similar in many ways. Getting your head around what could these things be doing and why – getting an intuition for that.
[00:16:38] Ilya Venger: The nice thing is that the more we use AI or LLM-based AI in our processes, the more intuition we get. But it’s not just about that. We can employ statistics on top of the responses we’re getting.
With a person, we usually don’t collect all the outputs. But with an agent that operates as part of the process, when we see the answer doesn’t conform to what we want, we can evaluate. We know the evaluation is not passing. What you can do is collect the logs – what were the inputs, what were the outputs? Then start analyzing that.
This is why we employ data scientists in our team – we’ve got 30 data scientists. They’re there, among other things, to understand what is going wrong. Then you need to start tweaking.
You either tweak the input data that you provide – the input data could be your RAG setup. Maybe something is going wrong because the retrieval from your documents is just not working. This happened to us so many times when we designed systems for financial analysis. If you get top 10 results from your documents, but your sorting algorithm is not very good, you need to resort and surface the right ones. Quite often this part is not working well. So the reason it fails is that it didn’t get the right information. We think it got the right information, but it actually didn’t retrieve it.
Second option is potentially I was not clear with my instructions to the AI in the way that AI understands it. So you tweak your system prompt, you tweak your context to make the system better.
And this alternative is usually under-explored but must be explored more: How should I be changing my process? Maybe I’m just not using it in the right place, or my process doesn’t require it in this place. Or I could restructure it such that the 80% is good enough, because this is exactly where I’m going to be introducing human in the loop.
With human in the loop reviewing the answers, there’s also a question: Is this going to be somebody who was supposed to be the next one in the chain of the process? Or maybe I can re-engineer it and say, if I’m working with customers, maybe with good UX I could push the customer to choose the right answer – present three answers and it would be clear for the customer which one is right.
There are more complex architectures currently starting to emerge where we’ve got AI checking AI. But I would not go there before I’ve explored simpler opportunities. Always start simple and then try to do things differently, and only afterwards complicate significantly.
[00:20:50] Noz Urbina: I’m very happy to hear all that. You mentioned system prompts – I want to jump in there because I agree it’s so important. System prompts and context were two terms you mentioned.
**[00:21:12] We’re doing a lot of work where we’ve developed a modular system for organizing how you give AI instructions. If anybody’s familiar with the DITA standard, it’s kind of like a DITA for AI. How do you set context for multiple AI who might be working together? You might be working with different engines, models, but you’re going to be defining roles within that – I am the data analyst, I am the financial analyst, I am the writer, etc.
What I’m seeing is a lot of users treat the AI as the magic genie – I will ask my question and my answer should come out. But when we’re doing this in business, the term system prompt refers to setting up instructions that define how a particular part of the system is going to work, and a key part of that is context.
When you say “act as a financial analyst” or “act as a writer,” you’re setting context. We do things like templates for saying: This is how you write a context. This is how you write an interaction framework. When someone says this, you do that. If I use quotation marks, this is what I mean – you should repeat this exactly to the user. If I use this kind of brackets, I mean this.
In the same way that you would be setting up a style guide or standard operating procedures with a team, you’re writing operating procedures and style guidance for your AI. This is where you fit in the role of other AI agents – if this agent sends you this, then do that.
Setting up the whole workflow and workflow design is where we went straight to at the beginning. The minute ChatGPT came out, we said the first thing we need to do is get away from just chatting with these things to defining workflows, defining roles, giving each role an interaction framework, context, what knowledge it’s supposed to have access to, what skills it’s supposed to have.
Then you can analyze each bit of the workflow in isolation. Is it failing at the writing stage? The data retrieval stage? You can explain it better to your human users – when you want this kind of task, go to this virtual agent.
You can do all of that before you get into the complications of agents checking agents. This is not a technical thing – you can do this with any out-of-the-box model.
What I’m seeing is a disconnect in the marketing from vendors. They are selling the genie instead of selling the business process approach. This is an amazing opportunity to re-engineer processes with human agents and artificial agents working together in re-engineered workflows. That is such a compelling pitch for business. But what I’m seeing is “this is a magic genie, ask it whatever you want and wonderful stuff will come out.”
[00:25:11] Ilya Venger: I’m not in sales and marketing – we’re actually building the genie. But you’re absolutely right. It’s interesting dynamics in the market – the market is moving very quickly. How we’re working with our product is evolving. We all collectively are starting to understand better what works and what doesn’t work.
It takes time. I don’t blame the salespeople or people in the field. Specialized consultancies are probably bridging some of that gap. Their job eventually is to sell. Hopefully you don’t leave your customer with “here’s some nice shiny toys, now go ahead and implement.” But quite often that is what’s happening in any software, whether it’s advanced software or otherwise.
[00:26:32] Noz Urbina: In their defense, this isn’t new to generative AI. I’ve been in this industry for a quarter of a century, and every new software package is always the solution to all your problems. That’s basically the marketing pitch for every software sale I’ve ever seen.
It’s not like they’re going in with this perspective and customers are resistant and skeptical. They’re going “Fantastic! Finally, the magic wand I’ve been waiting for.” The AI hype cycle is just regular tech hype scaled up. It’s the exact same thing.
Let’s go into when and how to go deeper. We’ve addressed that we need to think about the workflow and how we instruct more carefully to build processes. When should we consider AI checking AI? Or when should a business consider training their own model versus an off-the-shelf solution?
[00:27:59] Ilya Venger: When things are continuously failing and you’ve realized that just tweaking things on the level of one agent doesn’t work, and re-engineering your process completely doesn’t work – you first need to go through these motions.
For some reason, maybe from the early days two or three years ago when people started fine-tuning, I hear so much from different executives and even techie people saying “Let’s just fine-tune the model. It’s going to be great because we’ve got some data.”
Fine-tuning is one of the ways to affect models. For those that don’t know, it’s kind of retraining or additional training on top of an existing base model. This sometimes can yield very good results. It’s actually pretty difficult to do – it’s easy to throw training data at it, but you need good evaluations to understand whether this actually resulted in better results on your processes. It’s not very cheap because usually the fine-tuned models are more expensive to use afterwards.
For some reason it’s a go-to for everybody because people feel if they fine-tune the model, they’ve got a unique piece of IP in their hands. “OpenAI has a model, now we’ve got a model that is OpenAI’s model plus some delta that is our particular model.”
But actually your unique IP could be in other places. Like you said, setting up a really strong system prompt for your agent is actually very important. The IP can lie in different places, not just in fine-tuning.
Coming back to your question – when do we go there? When other modes have failed, then we go to further complication. I think with checking agents with AI, I would go there right now. The frameworks have matured and are continuously maturing. These are relatively simple engineering frameworks – I would do that before I would touch on fine-tuning.
When does fine-tuning make sense? In our group, we have fine-tuned quite a few different models for different goals. The reason to fine-tune is usually to find a better place on the trade-off between cost and quality.
[00:31:17] Noz Urbina: What’s the cost difference of fine-tuning?
[00:31:21] Ilya Venger: Let’s cover which models we’ve got in the market. I would separate models into large language models and small language models. Large language models have hundreds of billions of parameters. Small language models typically have up to 12-15 billion parameters, which means they require significantly less compute to run. They are cheaper to run, require less memory because they’re smaller, so they can fit in a smaller machine. They are easier and cheaper to fine-tune than large models.
Small language models are not as accurate, they’re usually not as “smart” as large language models, and they usually cannot reason in the same way. If we want to significantly improve a small language model, one approach is to add more data and continuously fine-tune it. Then for the price of running a small language model, you get something specialized to your use case. If you’ve added proprietary information from your organization, you might land in a better place on the trade-off between cost. Any request to that model would not cost one cent but 0.1 cent or even less. That is sometimes a competitive advantage.
But I would strongly caution before going in this direction. The advantage is not durable – the world is moving all the time. New small language models are coming out. You’re better off waiting for the next one, and then you’re going to need to fine-tune that one anyway. Once you fine-tune a model, it’s not done. If you fine-tune a small language model, you’re going to need to fine-tune the next generation of that model.
[00:34:12] Noz Urbina: So you can’t just upgrade the underlying model once you fine-tune?
[00:34:15] Ilya Venger: No, you can’t. There’s something called LoRA – there are ways to fine-tune, but these are never transferable between one model and another model. You can’t just take those adjusted weights and transfer them to another model.
[00:34:37] Noz Urbina: It’s like heavily customizing a piece of software. If you over-customize, you’ve locked yourself out of the future of that.
[00:34:49] Ilya Venger: Exactly. You would need to re-customize and re-fine-tune the new model. Potentially you might be able to use the same script, but that means every three months you’re going to need a team of data scientists doing that – running the script, running evaluations, checking whether the model works correctly. And the architectures of the underlying models are also changing all the time. So actually the training is usually not just lift and shift.
Second thing to consider is that the price of large language models is dropping continuously. We had a case – we trained a model last July, and then the large language model (GPT-4 in that case) – the price dropped almost tenfold. So the price differential we used to have where the small language model was cheaper than GPT-4, suddenly GPT-4 is 10 times cheaper and performing at the same level. The large language model costs essentially the same as my small language model.
Third reason is that serving small language models is a chore. It’s usually not just hitting an API. You need to think about how you’re going to host it, how many GPUs you need, where to get your GPUs – we know there’s quite a shortage.
[00:36:55] Noz Urbina: The technology infrastructure – you’re buying an elephant is what I’m hearing. You’re getting yourself into a lot of work.
I think you’re also referring to something I see all the time. For example, we had the “six finger problem” with AI-generated photos. I work a lot with the content industry, and many content people were quick to laugh off AI output because it was generic – the writing equivalent of the six finger problem. Either wildly inaccurate or terrible writing, or both.
But humans are not cognitively structured to deal with things that happen on an exponential curve. The difference in 18 months is so drastic, so night and day, that it’s hard for us to believe. We look at something and go “that’s blah, that’s how it is.” Then you go back to work, a year goes by, you turn around and the world has changed.
Riding that rocket is a whole mental model shift of the speed of change. I don’t know how we’re going to adapt to that.
[00:38:27] Ilya Venger: Your example about six fingers is actually very relevant because one of the things people have done is fine-tune models to get rid of the six finger problem. 18 months ago, some people invested the time in making the six fingers disappear.
[00:39:04] Noz Urbina: The six finger problem – when image generators first came out, they couldn’t get hands right. We say the six finger problem, but it could be the three finger problem or the 11 finger problem. We just couldn’t get hands to come out right when generating pictures of humans.
When you create a picture of a bee or tiger, maybe you didn’t notice. I remember it used to generate pictures of insects all the time and they looked great to me because I don’t know how many little things are supposed to be on their heads anyway. But with human hands and faces, we’re extremely sensitive. Teeth, eyes, fingers were very difficult for machines to get exactly right.
Then it got faces perfect because there are trillions of faces on the internet, but not actually that many good, useful examples of what hands are supposed to look like. It was a tough problem and took a couple years, which is a very long time in AI time. But from public popularity to really perfecting human hands did take less than 24 months.
Everyone was giving conference presentations going “Haha, look at these things they call intelligent, they can’t even get a hand right.” What you’re talking about – investing in a model, thinking this model costs X and there’s an ROI – by the time you’re finished and roll it out trying to get some ROI, your whole business case is gone because the underlying problem you were trying to fix 12 months ago has just vaporized.
[00:40:48] Ilya Venger: Yes, correct. Just one last piece on this fine-tuning problem. When does it work and when should you invest? You should invest in places where you need something extremely specific. The more specific the task, the easier it will be to evaluate, the easier it is to find the right data rather than just throwing everything at it.
This particular task could be, for example, named entity recognition. Today’s models are great at named entity recognition – we’ve had that for 15-20 years. Named entity recognition is recognizing concepts within a document or text – recognizing the name of a place and tagging “this is Madrid, this is London, this is Tel Aviv” as cities, or recognizing countries.
Big models are pretty good at this, but you can train a small model that’s going to be amazing at it for your particular set of entities. If your entities don’t exist much on the internet – like “I need to recognize the names of machines on the manufacturing floor, and I need to know whether this is a machine on the manufacturing floor or a brand of sneakers I’m wearing” – you could train a small language model to discern between the two very easily and become a specialist in this.
If you get to a problem where some particular task is failing, where you have proprietary data, and it’s a very particular thing that you want to fix that is not generic, then yes, that might be where you get good ROI on fine-tuning.
[00:43:11] Noz Urbina: I have to address that example because I know some of my audience will be thinking we were doing that with what we call auto-taggers far before the generative AI movement. Why would I use generative AI to do that when other machine learning tools – we’ve been doing auto-tagging with taxonomy and graph tools for easily a decade. Can you give another example of where fine-tuning generative AI might be good?
[00:43:57] Ilya Venger: To answer your first question – it’s not necessarily the best tool, but it’s very versatile. You don’t always need to train for a particular entity instruction. It’s actually very easy and quite often the results are significantly better than the older generation. They’re more expensive, but significantly better results.
Where else fine-tuning was useful – in recognition. One project we’ve done is recognition in agriculture, recognizing crop diseases. That was an image processing model we fine-tuned to recognize different diseases so a farmer could take a photo of a crop in their field and get the response “it’s being afflicted with this parasite or fungus, therefore you need this chemical to treat it.”
These things aren’t wildly available on the internet. It’s proprietary data to agricultural companies that have enough training data. The task is very specific – I need to recognize one of 150 diseases. When the differences between them are quite subtle, you could train a specialist model for this.
[00:45:48] Noz Urbina: Now you got me. For example, an auto-tagger could tell you “this support call log is talking about these products,” but a large language model would be able to read through many call logs. If you had trained it on your devices’ particular faults, common failure modes, etc. – you’re the only manufacturer who makes your product, the only one who can take photos of what the engine looks like when it’s leaking.

