Data Corpus: Why Smallholder AI Starts With What We Don’t Collect

AI is only as powerful as the data behind it—but what happens when that data is fragmented, incomplete, or excludes the very farmers it’s meant to serve?

In this episode of Grounded Intelligence, David Bergvinson is joined by Soumya Alamaru, Senior Consultant at Athena Infonomics, Jawoo Koo, Senior Research Fellow at the International Food Policy Research Institute (IFPRI), and Michael Minkoff, Independent Consultant, Former Director of AI and DPI Services at Athena, to explore why the future of agricultural AI depends on representative, locally grounded data. They discuss fragmented data systems, tenant and women farmers missing from official records, shared agricultural data corpora, data governance, and how localized AI models can deliver more accurate, trusted, and actionable advice for smallholder farmers. Through examples from Andhra Pradesh, including AP AIMS flood alerts and farmer-level data collection, the conversation highlights how better data can transform AI from generic recommendations into personalized agricultural decision support.

Credits

Music by Shuko Musemangezhi, Principal Advisor, Dev-Afrique

Transcript

[00:00:00] David Bergvinson: When we talk about large language models, we really need data to, you know, drive their development and delivery of valued services. 

 

[00:00:07] Soumya Alamaru: AI in agriculture is only as good as the data it is built on. Data exists, but it sits in silos. We don’t have a data scarcity problem. We have a data fragmentation problem.

 

[00:00:18] Michael Minkoff: When you are trying to drill down into who should make these decisions, how should they be made, it’s gotta be really contextualized to the realities in that country or subnational context in order to fit the right kind of data governance and data sharing. 

 

[00:00:32] David Bergvinson: How do we generate the local data set so that we, uh, are not actually delivering harmful information that in the context of the environment we’re trying to serve?

 

[00:00:43] Jawoo Koo: So I think, but the real value, um, you know, from researcher perspective is more strategic decision and taking calculated risk so that, uh, we can improve farmers’ livelihood and benefit and income, uh, at the end of the day and the end of season.[00:01:00] 

 

[00:01:08] David Bergvinson: My name’s David Bergvinson, and welcome to Grounded Intelligence. This is a series that looks at the responsible use of AI by engaging the community on their perspectives on the use and application, and some of the concerns around artificial intelligence to empower farmers to realize their full economic potential, and to manage the, the challenges of small scale farming in emerging markets.

 

[00:01:31] David Bergvinson: And so, uh, we’ll be going through, uh, a whole series on this topic, starting with, uh, data corpus or data. And I’m joined by, uh, three very insightful guests who are dealing with fragmented data systems and transposing those into valuable integrated data sets to drive models, to generate agricultural insights that help farmers make better decisions on the ground.

 

[00:01:57] David Bergvinson: Some of the terms you’re gonna hear are, like, frontier [00:02:00] models. These are large language models that have been developed by large companies like Microsoft or OpenAI, and, uh, you may be familiar with some of their names, like ChatGPT or Claude. Um, but we’re also gonna be talking about small language models and the data that, uh, is needed to drive those small language models, which is a critical gap.

 

[00:02:19] David Bergvinson: So we’re gonna explore these in the discussion ahead. So we welcome your suggestions. Uh, please leave those below on, uh, questions that have been raised through this conversation, and suggestions of future conversations as we work together towards unlocking the power and potential of artificial intelligence to enable farmers to realize their full economic potential.

 

[00:02:48] David Bergvinson: So thank you for joining us on Grounded Intelligence, and this is part of a broader community called, uh, AgX AI, where we build, um, [00:03:00] really a community around the responsible use of artificial intelligence to empower farmers. And so if that topic is something that interests you or you wanna contribute, I encourage you to subscribe to this channel so that you don’t miss upcoming podcasts and interviews.

 

[00:03:14] David Bergvinson: But also, uh, go over to agxai, uh, .com, where you will find more information around community events, uh, whitepapers or discussion papers on key topics related to responsible use of AI, and help each other, uh, find opportunities to secure resources and collaborate on projects in order to unlock the power of AI to empower farmers to realize their full potential.

 

[00:03:40] David Bergvinson: So we look forward to your con- contributions. Uh, please leave comments in the, uh, the comment section and subscribe to track our progress and contribute to the ongoing conversation. Until next time, thank you.[00:04:00] 

 

[00:04:00] David Bergvinson: Well, uh, so welcome everyone. Uh, today we’re gonna be looking at the topic of data. And so in the series of Grounded Intelligence, there’s many aspects to the responsible use of AI, and data is a core component, uh, because without data, we can’t really drive the whole system. So this is a foundational topic, and we’re privileged to be joined by three experts.

 

[00:04:23] David Bergvinson: Um, Jaewoo, uh, from IFPRI, and Soumya and Mike from Athena Informatics. And so I’m gonna let each of them introduce themselves just so that you can understand the experts that we’re, we have, uh, to facilitate this discussion. So Jaewoo, would you mind a brief introduction? 

 

[00:04:43] Jawoo Koo: Yeah, sure. Yeah, thanks. Yeah, hi, everyone.

 

[00:04:45] Jawoo Koo: My name is Jaewoo Koo. I’m a senior research fellow at International Food Policy Research Institute based in Washington, DC. Uh, my background is crop modeling in agriculture engineering field. So I work on many different types of [00:05:00] data, uh, that goes into the crop model, and also trying to understand, um, you know, what models rather generate.

 

[00:05:08] Jawoo Koo: So yeah, it’s very interdisciplinary research I’ve been doing at IFPRI, uh, for over almost two decades now. 

 

[00:05:15] David Bergvinson: Thank you, Jaewoo. Uh, Mike, over to you. 

 

[00:05:19] Michael Minkoff: Sure. Uh, thank you, uh, David. So, uh, Mike Minkoff, and for the… I’ve been overseeing, uh, operating as the director of artificial intelligence and digital public infrastructure services at Athena.

 

[00:05:33] Michael Minkoff: Um, uh, and been particularly focused in this kinda ag, uh, XAI community of practice and, and working on, uh, issues around especially, uh, data corpus and benchmarking, uh, deeply in that community of practice for the last year on the heels of a 15-year career in global development, um, that spanned working in environment and sustainability, climate risk, [00:06:00] and, uh, more recently, kinda data systems, data governance, and emerging technologies.

 

[00:06:07] David Bergvinson: Thank you. And Soumya, please. 

 

[00:06:11] Soumya Alamaru: I am Sowmya Alamuru. I am a Senior Consultant in Athena Economics. I did a PhD in economics. Now working closely with AP government, we are providing technical assistance to Andhra Pradesh, which is a state in India, on AI advisory. I’m leading the project from Andhra Pradesh.

 

[00:06:34] David Bergvinson: Wonderful. So we have some practical field experience, so we’re well represented here. Um, so our opening topic, as I said, you know, around data, w-when we talk about large language models, we really need data to, you know, drive their development and delivery of valued services. And so let’s talk about why data comes first.

 

[00:06:54] David Bergvinson: So, Jau, would you, uh, you know, give us a little bit of a tour around the AI models and [00:07:00] tools, uh, that we’ll be exploring. And from your perspective, why does the data layer matter so much in realizing the benefits of AI in agriculture? 

 

[00:07:09] Jawoo Koo: Yeah, sure. So yeah, again, because my background is crop modeling, I always have that, uh, tendency to apply all these data topic and similar discussions in the context of crop modeling.

 

[00:07:21] Jawoo Koo: So bear with me. Um, so crop models, for example, uh, it’s not exactly AI model, uh, but it has the same constraint on data. So any model, including AI models are, are only good, as good as the data it, uh, we may provide, uh, to develop the models for. So for example, uh, if the models are only trained or developed based on specific types of environment, say like North American agricultural systems and maize farming in large scale farming, uh, scale farms, uh, it may not have all, uh, [00:08:00] good enough understanding of what are the challenges are farmers facing, how to address the risk, and what kind of advices, uh, that it needs to provide under certain climate conditions, uh, for small scale farmers in developing countries where we, we are focusing on.

 

[00:08:16] Jawoo Koo: So that will introduce really inherent biases. Uh, the models might start giving advisories that are not relevant for smaller farmers, not really actionable. It might provide advisories to access tools and use chemicals that are not even available in small holders, uh, farming systems in the countries where farmers are part of.

 

[00:08:38] Jawoo Koo: So the, the kind of, uh, the limitations, uh, uh, all start stemming from the data. So the data we provide to the model and the models are trained for really need to be, uh, extensively Uh, com-comprehensive for the environment that where of our target our users are, are, you know, [00:09:00] you know, cultivating farms, cultivating crops for.

 

[00:09:03] Jawoo Koo: So that, that really worries me. Uh, so all, all these large scale, uh, hyperscalers, hyperscaling models are, um, you know, they’re using data from English language in well-defined, uh, kind of farming systems and well-documented systems. So it may not ha- capture all these, uh, nuances or, or, or really detailed informations, uh, on smallholder farming system.

 

[00:09:30] Jawoo Koo: So yeah, it, it might miss all these, um, details or the challenges of smallholder farmers. 

 

[00:09:36] David Bergvinson: Yeah, the context of that data is very important, as you point out. Yeah. So, um, so as we dig deeper into the understanding the problem of data, uh, I’m gonna turn to Mike and Soumya on this, um, because a lot of the current AI efforts in emerging markets relies on existing datasets.

 

[00:09:55] David Bergvinson: So Soumya, first to you, uh, drawing on your experience in Andhra Pradesh, India, in scaling [00:10:00] AI advisory, what do you see as the biggest challenges around data access to drive, uh, AI-enabled advisory services? 

 

[00:10:08] Soumya Alamaru: Well, as Jau correctly said, AI in agriculture is only as good as the data it is built on. And in regions like Andhra Pradesh, the biggest challenge is not lack of AI, but the lack of reliable representative and usable data The data is fragmented across departments.

 

[00:10:29] Soumya Alamaru: Data exists, but it sits in silos. We don’t have a data scarcity problem, we have a data fragmentation problem. For example, agriculture department holds crop data, farmer level data. AP, AP Planning Department manages procurement data. AP Space Application ca- holds satellite, uh, s- um, satellite imagery, and banks hold credit data.

 

[00:10:53] Soumya Alamaru: None of these were designed to talk to each other. Moreover, institutional silos lead to duplicated data collection [00:11:00] and incompatible formats. Another challenge is last mile data gaps. The most granular actionable data, what a farmer… Example, what a farmer in one of the districts actually planted, when they irrigated, what inputs they used, rarely gets captured digitally.

 

[00:11:22] Soumya Alamaru: Village level agricultural assistants still work on paper in many corners of Andhra Pradesh. 

 

[00:11:31] Michael Minkoff: Mm-hmm. 

 

[00:11:32] Soumya Alamaru: The third, perhaps most critical and deeply structural challenge is the mismatch between farmer identity and land records. In AP, more than 80% of cultivable land is operated by tenant farmers. Among those, a substantial portion is informal tenant, uh, farmers.

 

[00:11:54] Soumya Alamaru: They often don’t appear in official datasets These are the [00:12:00] major challenges that are there in Andhra Pradesh. So what you get is AI system trained on skewed data if you don’t consider these challenges. 

 

[00:12:10] David Bergvinson: Yeah, good point. Uh, Mike, you know, so what is the big risk then, given what Ramya just said, around building these systems that are using data that really isn’t designed to serve small-scale producers in mind?

 

[00:12:25] Michael Minkoff: Yeah, I mean, and thanks, Sowmya, for that really, I think, uh, eloquent and vivid uh, articulation of what this really looks like on the ground, um, in, in the con- certainly, at least in the context of Andhra Pradesh. And, and, you know, I think i- in our experience, Andhra Pradesh is, is pretty far advanced relative to a lot of other, uh, low and middle income contexts.

 

[00:12:45] Michael Minkoff: So, so I think that’s important to note as well when you’re thinking of this broader question of what these risks are. And the risks are you’re, you’re gonna be, you know, producing and putting out advisory that may not [00:13:00] be really particularly relevant or accu-accurate or contextualized to the needs of the actual individuals that are receiving that AI-generated advisory.

 

[00:13:08] Michael Minkoff: And a lot of times that’s gonna be based on hyper-local, um, credit information, ac-access to credit information, or soil condition information or, or microclimate or, you know, weather information. Um, and, and you have these kind of structural gaps and, and, you know, possibly in data, possibly in data integration, possibly, um, in just representativeness, uh, that, that need to be understood and acknowledged as you’re thinking about how to build out, uh, really meaningful and effective, scalable, sustainable, uh, AI advisory solutions and approaches.

 

[00:13:45] David Bergvinson: Great. So you mentioned the data gap. So Sowmya, in your experience, what are the farmer realities that are most frequently, uh, not considered or excluded or distorted in the current data ecosystem, drawing again on your AP work? [00:14:00] 

 

[00:14:01] Soumya Alamaru: As I said earlier, tenant farmers- Mm-hmm … they are out of the data structure.

 

[00:14:07] Soumya Alamaru: AP has a significant portion of cultivated land is, is operated by tenant farmers who are invisible to government data systems. They don’t appear in land records, don’t receive crop insurance or schemes. Uh, their use of inputs, yields are not captured anywhere. This is one of the, uh, biggest issues. What percentage 

 

[00:14:28] David Bergvinson: of the farmers are tenant farmers?

 

[00:14:32] Soumya Alamaru: Sorry, com- 

 

[00:14:33] David Bergvinson: What percentage of the farmers in AP are tenant farmers? 

 

[00:14:36] Soumya Alamaru: Uh, percentages-wise, it is, like, around thirty to forty percent, but eighty percent of cultivable land is operated by tenant pers- uh, tenant farmers. 

 

[00:14:48] David Bergvinson: So the, but those- They- … farms would have data, though. It’s just that they don’t have data on the tenant farmer who’s cultivating that land.

 

[00:14:56] Soumya Alamaru: Yeah. The, the, the, like, the kind of crops [00:15:00] they sow or the inputs they use or the market, um, or the yields they, they arrive at are not captured anywhere. And most of these farmers hold, uh, most of these farmers, uh, would be cultivating in less than two acres of land. And these government data collection thresholds have, uh, government data collection systems have thresholds like, you know, uh, less than one acre or two acres are not captured in the data systems.

 

[00:15:32] David Bergvinson: So that’s a big challenge. Um- Yeah, 

 

[00:15:33] Soumya Alamaru: it is. It is … I know you’re 

 

[00:15:34] David Bergvinson: working on. Yeah. 

 

[00:15:35] Soumya Alamaru: Yeah. 

 

[00:15:36] David Bergvinson: Um, 

 

[00:15:36] Soumya Alamaru: I, I- Second is l- land is disproportionately registered on, in men’s name. In Andhra Pradesh or in India, farming is a family business. It’s not, it’s never one person’s job. The whole family, husband and wife and their children in, gets involved in the farming In some ca– in most of the cases, probably women [00:16:00] take farming decisions, but the lands are not their names.

 

[00:16:05] Soumya Alamaru: Mm-hmm. We don’t consider women farmers as farmers, and the data doesn’t capture that part of it. 

 

[00:16:12] David Bergvinson: Yeah. You’ve raised some two very good examples around inclusion, and we’ll, we’ll circle back to those, uh, later. You know, one of our other big challenges, as you said, Somya, is the fragmented data. And so, you know, going to this topic of, uh, shared data, we call it a corpus or a body of data.

 

[00:16:32] David Bergvinson: Uh, Jawoo, I’d like to get your perspectives on, um, the discussion papers that AgX AI is working on, especially as we talk about data corpus and the sharing of data. Um, why is this concept so important, um, in addressing this isolated data ecosystem that agriculture faces? 

 

[00:16:52] Jawoo Koo: Right. Uh, so there are two things, um, for, for this question.

 

[00:16:56] Jawoo Koo: So one is, uh, to really understand what kind of language [00:17:00] the farmers are speaking, even in just agriculture in general, uh, which may or may not be similar to what general purpose like large language models are trained on. Agriculture languages and, um, norms and specific terminology, some are really colloquial terms that only kinda used in, um, you know, farming context or again, like may or may not be so central to the large language model.

 

[00:17:25] Jawoo Koo: So the providing and managing and consolidating this agriculture data corpus is really important if we are serious about, uh, providing real actionable and, um, you know, meaningful advisory for farmers. Uh, they wouldn’t take advisory seriously, uh, if we use very generalized language, not really speaking the same tone or same kind of, uh, the terminology.

 

[00:17:52] Jawoo Koo: So yeah, I think that’s really important to communicate better, uh, with farmers and understand better what kind of language they are speaking. Also, in [00:18:00] generally, uh, this kind of corpus sharing, shared data corpus provide, um, again, the more broader context to the types of, uh, information and advisory we are speak- uh, providing.

 

[00:18:12] Jawoo Koo: It’s not just agriculture, like on the ground and field levels expertise and field level practices. It, it also needs to touch on the labor, agriculture labor and financial access, and everything in rural community also need to be captured. And again, this just cannot be done, cannot be captured in one single or limited number of data sources.

 

[00:18:33] Jawoo Koo: So this shared data corpus is really important for- For reaching out to farmers and providing meaningful advisories. 

 

[00:18:41] David Bergvinson: Yeah. So you, you lead to an interesting point here, um, really around the balance between scale and specificity of these datasets. Soumya’s already mentioned that, you know, s- large segments of society are missing in these representative datasets.

 

[00:18:55] David Bergvinson: So, you know, how do we find the right balance? Because going into that very [00:19:00] granular data, that, that costs quite a bit of money to collect. So how do we find this balance between the specificity, uh, and scaling these, these geographically large datasets and, uh, at least from the CG perspective? 

 

[00:19:14] Jawoo Koo: Yeah, I, I know.

 

[00:19:15] Jawoo Koo: So this is really challenging, uh, in a way . As a research institution, things are rapidly changing so fast. Um, so yeah, it’s, it’s even just hard to just k- keep, um, keep track of what’s going on, uh, where we should focus. So the getting the right balance is absolutely critical. Uh, I think the general agricultural knowledge, um, and, and to some extent the country level, large scale, large economy level information and data and knowledge are very well captured in hyperscalers and a global scale, uh, corpus kind of exercise already.

 

[00:19:51] Jawoo Koo: So I think what we can do, again, as a institution like CGIAR, um, is to provide more detailed, focused, um, [00:20:00] very thoughtful, uh, capturing of the processes that farmers are, you know, um, challenges that farmers are dealing with. Like for example, uh, if we are to provide advisory for like, say, for chili farmers in India, the general farming practices or the best practice for chili farming in India must be already captured very well in the large corpus or, or the existing corpus already out there.

 

[00:20:27] Jawoo Koo: But what we can do is what is the latest trend of, of, of the managing chilies? Where is the market demand or emerging kind of consumer demand? And some-something more, um, actionable, more dynamic, um, types of data. I think that will be probably more meaningful, uh, effort that we can provide so that we can be complementary, uh, and to reach the balance of the specificity and scaling.

 

[00:20:53] David Bergvinson: So that’s a great segue into some broader issues around go- governance, incentives and trust, which is so [00:21:00] important. Um, so if you betray the trust of the consumer in this, it’ll take a long time for them to come back to this technology. So Soumya, I’d like to ask you around data governance. Um, it’s a term that feels abstract, so why don’t you give it some context as it relates to the reality on the ground in AP around data sharing, access, and the responsible use of personal information?

 

[00:21:25] Soumya Alamaru: In AP it’s just the AI advisory s- is it resounding, my voice? 

 

[00:21:35] David Bergvinson: Yeah, it’s okay. We can hear you. 

 

[00:21:37] Soumya Alamaru: Okay. So I can talk about, um, A- AP j- AI advisory is still growing in, in the sense like we really don’t have advisory systems that, uh, actually considers granular level data or, uh, use a data corpus which is, which can, uh, [00:22:00] which can, uh, includes granular level to aggregate data.

 

[00:22:04] Soumya Alamaru: But what a representative data would be looking like is to, uh, cover all agroeconomic zones in AP and major eight to 10 cropping systems and irrigation systems like rain fed, canal or borewell or irriga- or, or tank irrigated context, and full spectrum of farm size from .5 acres to 20 plus acres. If a dataset covers all of these aspects, we can say that is a representative data.

 

[00:22:42] Soumya Alamaru: And if I have to give a example, recent example of a good AI model from Andhra Pradesh, a model, uh, that uses Godavari Delta region, um, it’s a real time Goda- Godavari is a river in Andhra Pradesh. It is real [00:23:00] time Godavari flood advisory is a game changer. When the state integrated IMD weather feeds with canal flow data and farmer location data, it became possible to push geo-targeted text-based flood warnings to farmers in specific regions.

 

[00:23:16] Soumya Alamaru: These flood alerts help farmers in many ways, not just the safety point, but to make a decision on sowing or harvesting

 

[00:23:28] Soumya Alamaru: And recently, like last year, like a few, uh, another example I can give of– give is, uh, ahead of cyclone month into 2025. AP AIMS, which is a, our data re-re, uh, repository system that recently Andhra Pradesh government has started building in, delivered 3D flood simulations and early warnings via SMS and IVR to millions of farmers, which helped in protecting eight hundred thousand [00:24:00] hectares from flood damage 

 

[00:24:04] David Bergvinson: Mm-hmm.

 

[00:24:05] David Bergvinson: So, so I mean, that’s really important, especially in the, against the backdrop of, backdrop of climate change and extreme rainfall events. Um, I, I, I, I turn to you, Mike, around, okay, at a higher level, who’s making the decisions and on what data is made publicly available, and how is it regulated? Uh, maybe just drawing on your experience again with AP or perhaps African nations that you’ve worked with 

 

[00:24:33] Michael Minkoff: Yeah.

 

[00:24:34] Michael Minkoff: I mean, I think, I think that the question of who is making those decisions or who should make those decisions is an interesting one and a challenging one. And I think, I think what we’ve seen, and I know, you know, in, in other fora, uh, Jiawu, we’ve kind of touched on this, right? Um, where there, there’s definitely not one size fits all.

 

[00:24:55] Michael Minkoff: Uh, and, you know, the models that are success- I mean, even if you look at [00:25:00] India, India’s got a, a variety of models in the way that they are approaching kind of the challenge around building data corpus or corpora. Um, there’s, there’s a kind of federated structure that’s been built at the federal level through Open AgriNet and, uh, AgriStack and Vista, you know, now Bharat Vistar.

 

[00:25:20] Michael Minkoff: And then there’s kind of state level deployments or adjacent deployments. Andhra Pradesh is a kind of an adjacent deployment to that. Uh, Ma-Maharashtra has a state level. And all of these have kind of different function, you know, different features and different go- roles of government. And then if you look at, you know, contexts in, in certain countries in sub-Saharan Africa, uh, it’s very different as far as what is the existing digital infrastructure capacity and capability, what is the existing government kind of capacity and levels of trust in government to fund and monitor and manage something like this effectively.

 

[00:25:54] Michael Minkoff: So the, that’s all a, a, a kind of long preface to say, [00:26:00] I think when you are trying to drill down into who should make these decisions, how should they be made, it really has to be, um, contextualized to the political economic reality, the, the techno institutional realities, uh, the, you know, the sociocultural realities.

 

[00:26:17] Michael Minkoff: Choose your hyphenate, uh, uh, hyphenate terminology. But it’s got to be really contextualized to the realities in that country or subnational context, uh, in order to, to, to fit the right kind of data governance and data sharing, um, processes and protocols, uh, that, that can really kind of sustain. Um, and, and those need to be designed with real clear eye on kind of the what’s in it for them for all of the parties involved, right?

 

[00:26:48] Michael Minkoff: And like, how, how are they gonna, you know, what are the meaningful incentives to sustain contribution and collaboration through this data sharing mechanism and, and partnership? And so whether that’s a, a [00:27:00] government entity that’s contributing, whether it’s a research institute l- or like network like CGIAR, whether it’s private sector actors that you, that have great data but have competitive interests and maybe not sharing all of it, uh, you know, all, all of them have their different- 

 

[00:27:16] David Bergvinson: Sure

 

[00:27:16] Michael Minkoff: realities. Yeah. 

 

[00:27:18] David Bergvinson: Well, let’s, let’s, let’s pivot to Jiawu then from the research context perspective around the data sharing, uh, and the incentives for that. Jiawu, are there You know, incentives or policies in place that facilitate data sharing across the CG and make it accessible to a broad range of stakeholders in country 

 

[00:27:36] Jawoo Koo: Absolutely.

 

[00:27:37] Jawoo Koo: So CGIAR, as the largest publicly, uh, funded agriculture research institution in the world, uh, we do have a policy since 2021 on, uh, fair and open ac- open data access across all the CGIAR centers, including my center, IPRI. So all the research output, research data, research publication should be in the public domain, [00:28:00] um, um, eventually.

 

[00:28:01] Jawoo Koo: We, we are given a little bit of embargo period, uh, while we are developing and finalizing publication, but the idea is that everything publicly funded project, uh, that produce should be in the public domain. But having said that, we do have a challenge. We have a little bit of issue or, or say mismatch between the purpose of this policy and what’s really useful corpus for, like, say, uh, training AI and A- LLMs and AI models for advisory services.

 

[00:28:29] Jawoo Koo: The nature of our research is to develop and, and generate, uh, knowledge, knowledge around how we improve farming and how we improve productivity, et cetera, or this, um, you know, the toward good outcomes and better outcomes in food systems. Uh, this will may, may or may not be the reality in farmers’ field.

 

[00:28:50] Jawoo Koo: So the– we, we often have, like, say, uh, agronomic trials, very controlled environment there where we know what’s going in and what’s coming out. We [00:29:00] can measure perfectly, uh, we can, uh, observe perfectly so that we have better understanding of what works and what doesn’t. Um, so again, that this, this may or may not be really relevant for smallholder farming system.

 

[00:29:13] Jawoo Koo: So, uh, so we always have a little bit of, um, um, the issue, issues on trying to fully, uh, understand and digest the research output, research findings, and turn that into agronomic, uh, or, or farming advisory, uh, is always a little bit of, uh, kind of… There is a human really need to be in the loop, so they interpret, uh, that research finding into the right context for farmers.

 

[00:29:38] Jawoo Koo: So that’s a little bit of issue. Um, so having said that, research data is again, like, still, still fully open. We have a policy. Uh, we also have a governance framework, uh, we are developing that define four specific roles, uh, in this data ecosystem: data owner, data lead, data producer, and data users. And, uh, we are also [00:30:00] articulating whose responsibility at what stage, at what level, uh, so who will make such decision on data sharing and opening data or, or, um, yeah, in different, different ways.

 

[00:30:12] Jawoo Koo: So yeah, it’s, it’s a lot of very complex issues, and we are dealing with. But again, a research institution can do so much, so we really need partners on the ground to make sure what we are generating is, uh, fully useful and meaningful for farmer to 

 

[00:30:26] David Bergvinson: use. Yeah, and all this is happening in the context of things moving so fast in the AI world.

 

[00:30:31] David Bergvinson: You, you mentioned frontier models then for, for those Not familiar, these are, you know, the models that have been generated, uh, by Anthropic called Claude or through OpenAI called ChatGPT, um, Gemini in case of Google, uh, Meta, et cetera. And, uh, they’re moving at such a pace, and yet the clients that we’re trying to serve, namely smallholder farmers in emerging markets, uh, that’s a very different ecosystems.

 

[00:30:58] David Bergvinson: And so there’s some hard choices and [00:31:00] trade-offs to, to manage here. Um, so in light of that comment, uh, and these rapid- rapidly moving frontier models, how do we generate the local data set so that we, uh, are not actually delivering harmful information that isn’t in the context of the environment we’re trying to serve because it’s coming from, you know, uh, an advanced economy data set?

 

[00:31:25] David Bergvinson: Uh, how do we sort of round that off? And Mike, maybe just give some of your high-level perspectives on this given your India experience. 

 

[00:31:34] Michael Minkoff: Yeah, I mean, what’s really, you know, one of the interesting parts of this, right, it’s, it– A, there’s kind of a juxtaposition or trade-off between the use of the frontier models that you described just– and, and do they have the good data in them, period.

 

[00:31:49] Michael Minkoff: But even if they do, there’s a different question of do we wanna be using frontier models in these contexts all the time or, or… I mean, increasingly we’re seeing [00:32:00] how powerful and how accurate small, small language models, localized small language models can be, uh, if they have the right data in hand. And, and there’s also these kind of ongoing subtext questions around sovereignty of models and, you know, on-premise solutions and on-site data storage and management.

 

[00:32:21] Michael Minkoff: And so, so I think there’s, there’s a couple layers to it, and I think, you know, to collect that data, right, as small language models get more powerful, uh, you actually– it lo- it starts lowering the burden of how much good data you need to get to be effective, uh, to, to be able to develop effective localized models that, that can, uh, fill in some of these gaps that these frontier models don’t have.

 

[00:32:48] Michael Minkoff: And that also then helps you solve for the sovereignty challenge and on-premise data management challenges and compute cost and infrastructure cost challenges that are, that are real, [00:33:00] real limiting factors in a lot of low and middle income country contexts. So, so I would say, you know, continuing, you know, through programs like what we’re doing in Andhra Pradesh right now or other programs, that gives you really a, a opportunity to start collecting this localized data or, or work that CDIR can do and kind of different approaches to, to complementary data collection like Jiawu.

 

[00:33:25] Michael Minkoff: I- if I, if I was, uh, capturing what you were saying correctly, you know, that– those are– those present opportunities to start building these kind of complementary, more localized, you know, uh, data sets that can feed into whether it becomes a frontier model at a, a large language model scale or it becomes small language model deployments.

 

[00:33:46] Michael Minkoff: Um, I- either way, I think, I think– and, and from what I’m seeing, it seems like small language models will increasingly be the way to go. I think, I think that’s, that’s kind of where the opportunity, uh, continues to [00:34:00] emerge. 

 

[00:34:00] David Bergvinson: Well, and especially around the data sovereignty and model sovereignty that you mentioned, uh, that’s important.

 

[00:34:05] David Bergvinson: So, you know, I, I– and you brought this up earlier on, Jiawu, is, you know, we’re using that data to make decisions through modeling. Um- And so, you know, you also pointed out that we need quality data to drive a quality model that dri- delivers relevant results. Um, so what does good data look like to you that enables farmers advisories, um, or, uh, system-level interventions to be effective, uh, using this new technology?

 

[00:34:34] Jawoo Koo: Right. Uh, so, uh, through our partnership with, uh, Digital Green, um, who are already scaling, uh, this agricultural advisories, AI-powered agri- agricultural advisories to many, many farmers, uh, so we actually had a chance to look into what kind of questions farmers ask. A-and there are just two groups of, um, you know, two, I guess, a bucket of questions.

 

[00:34:58] Jawoo Koo: One are very strategic, one [00:35:00] are very tactical. So tactical questions like what can… What should I do right now, uh, for addressing my problem? What should I do today or tomo-tomorrow? And then there are more strategic questions. Okay, why should I, uh, how do I change my farming to be more profitable or more resilient to climate change and things like that?

 

[00:35:19] Jawoo Koo: Uh, what we saw was that AI models are currently used for very tactical questions only. And sometimes farmers even ask questions they already know answer to. They are just testing how good this model is. So I think but the real value, um, yeah, from researcher perspective is more strategic decisions and making calculated, uh, taking calculated risk so that, uh, we can improve farmers’ livelihood and benefit and income, uh, at the end of the day and the end of the season.

 

[00:35:48] Jawoo Koo: So I think more the more sophisticated, uh, system we are developing and more comprehensive data we are providing, I think we can enable farmer to think more, uh, medium to long term, um, [00:36:00] kind of g-goals and, and developing vision, then plan for that so that they can really improve their livelihood over longer term and more, yeah, thriving in farming.

 

[00:36:10] David Bergvinson: Yeah. Well, and I, and I think, uh, you know, in my experience, the market integration is a really important component of this. So it’s the advisory to produce the crop, but as you point out, thinking even beyond that strategically, what market is that crop gonna go into and planning accordingly. So very good point.

 

[00:36:28] David Bergvinson: Uh, Sawмы, I turn to you, um, around examples where data, uh, design and collection has materially changed the outcome of a project. So you started with a protocol for either pulling in existing data or even surveying for new data, but found that, uh, they weren’t meeting the needs or delivering the kind of benefits you anticipated.

 

[00:36:51] David Bergvinson: So, um, how, how has that played out, uh, in advisory services in AP where you’ve, you’ve seen this kind of situation where [00:37:00] data quality or access has really limited the ability to deliver the impact desired by the, the government, uh, the, the key stakeholders and farmers themselves? 

 

[00:37:12] Soumya Alamaru: Yeah. As I mentioned earlier, like, um, I just want to add one point to the– when we are talking about what kind of data do we have to collect and how, uh, we– how do we– on what basis we build, uh, AI systems.

 

[00:37:30] Soumya Alamaru: The data collection should take a bottom-up approach. Usually, what happens is, like, we have district level, state level data, and then we just use averages to, uh, proportionally distribute to lower levels, village level or farmer level. But if we start from the farmer level or village level, that what data exists, and then add up to, uh, you know, a mandal level or district level and, or state level, your data would be much [00:38:00] more close to the ground, and it would be reliable.

 

[00:38:03] Soumya Alamaru: So r-recently in Andhra Pradesh, uh, r-some real efforts have, uh, been put in place to make this data corpus. The government is working on building large integrated data repositories like AP AIMS, which is Andhra Pradesh Agriculture Information Management System, that aims to bring various data points, including cropping patterns, soil health, pest management, weather data, all into a single platform.

 

[00:38:36] Soumya Alamaru: So using this platform recently, uh, they have started, uh, delivering weather alerts. Uh, it could be like, uh, cyclone, uh, alerts or whirlwind alerts. These are really having an impact on farmer life. They, they– these are– these have impact on, like a positive impact on their farming [00:39:00] decisions, also on their safety nets.

 

[00:39:03] Soumya Alamaru: Re– That is one example I could given. And then if I have to talk about poor quality AI advice, which are like there are many AI a-agri fintech companies are releasing a lot of AI advisories which depend on standard university or, you know, uh, textbook guidance, which are, like, useless to particularly regional smallholder or, um, uh, you know, uh, or, uh, tribal area farmers.

 

[00:39:39] Soumya Alamaru: So those like the, the, the advisory has to be totally relied on ground root-ground-truthing or close to reality data. The, uh, I’m happy to say that in Andhra Pradesh, these efforts have been initiated, and they started giving good results [00:40:00] 

 

[00:40:02] David Bergvinson: Yeah. So, uh, and, and I guess you’re mapping these data gaps in the case of Andhra Pradesh and addressing them, um, and making sure they’re relevant, which is really encouraging to see, especially when you point out the issue of, um, you know, farmers that don’t hold title or women who are often, uh, out of the loop when it comes to formal extension services.

 

[00:40:23] David Bergvinson: And so this really opens up a huge opportunity for inclusion if it’s done properly. So, um, that leads to my- 

 

[00:40:31] Michael Minkoff: David, could I, could I just add one little bit on this, on the Andhra Pradesh? 

 

[00:40:35] David Bergvinson: Sure. 

 

[00:40:36] Michael Minkoff: Yeah, just very quickly. I mean, one thing that’s very cool about the, what the government of Andhra Pradesh is aiming to do and some of the work that we’re doing, it, it extends beyond, you know, agricultural data as we’re talking about it, and includes, and, and a big focus of, of the program, uh, we’re involved in is looking at kind of the horticultural data, the livestock data, the, the fisheries and [00:41:00] aquaculture data as well.

 

[00:41:01] Michael Minkoff: And it– the, the vision is to bring it all into an integrated space where all of this data can kind of talk to, to each other in addition to kind of market and, and kind of financial access information and data and climate and weather data. So you can get this really, you know, kind of Jawu, like what you were talking about this, this more sophisticated strategic inferential advisory.

 

[00:41:23] Michael Minkoff: Um, and, and so that’s kind of the direction it’s heading. And, and right now there’s a lot of data ecosystem gap mapping and, and hole filling and, and kinda integration that the team is doing. Um, but I, I just think that’s an important kind of dimension to, to highlight as well as the possibilities 

 

[00:41:41] David Bergvinson: that are there.

 

[00:41:42] David Bergvinson: Yeah. That’s leading into our, our last question, thanks, Mike, around our last two questions around looking forward. So what do you see in the future? Uh, I’m gonna ask you b- uh, frame both questions and we’d do a round table on your responses. The first one is, if we get the data corpus right, this federated, responsibly assembled, [00:42:00] contextually relevant data set What becomes possible for AI in agriculture, uh, that isn’t possible today?

 

[00:42:07] David Bergvinson: And the second question is, what worries you most if data remains fragmented? Um, so those two questions as we close off on this topic of data corpus. So let’s turn to Jaewoo first. 

 

[00:42:21] Jawoo Koo: Yeah, sure. So, um, uh, it’s, it’s not totally correct to say it’s not possible today, right? Because some things already happening, uh, but it’s hyper-personalized assistant, um, almost like persistently with you so to, to make, uh, suggestions, to make recommendation just for you.

 

[00:42:40] Jawoo Koo: I think that, that, that will be really, um, you know, transformative, uh, to some extent for smallholder farmers to understand what kind of resources you have, what kind of challenges you have, what’s the best course of action that minimize risk, et cetera. Uh, right now, I think when you start chatting session with any [00:43:00] chatbot, it kind of, uh, start from scratch or have only very limited session information, only focus on what you provided for the chatbot.

 

[00:43:09] Jawoo Koo: But, uh, if all these data, uh, sources are connected as a, like more, uh, digital public infrastructure, it can pull m- many different types of data from different sources, so it can provide more customized and really personalized a- advice for- 

 

[00:43:25] David Bergvinson: Yeah. Personalized advisor for farmers then. 

 

[00:43:27] Jawoo Koo: Yeah. Exactly. Yeah.

 

[00:43:29] Jawoo Koo: Okay. Um, and I, I think it can, uh, eventually lead to something more physical. I, I think there are, like, a lot of innovation happening in robotics. Uh, there, there will again, um, in a way LLM, yeah, um, so we, we have been hearing LLM is going beyond language, already past the language barrier. It’s just becoming multi-modal, uh, becoming o- operating system for robotics as well.

 

[00:43:54] Jawoo Koo: So I think it can really, uh, pave way for more, uh, physical and, [00:44:00] um, different types of AI transformation- Yeah … that we can lead to. Well, 

 

[00:44:04] David Bergvinson: I’m already seeing like, you know, drone integration of AI- Yeah … and analytics, for example, today. 

 

[00:44:08] Jawoo Koo: Exactly. 

 

[00:44:09] David Bergvinson: Yeah. It’s a, a hint of what’s to come. Um, Mike, over to you. Uh, your, your thoughts on those two questions.

 

[00:44:17] Michael Minkoff: Yeah, I mean, I, I largely agree with Jiaguo as far as, like, if we get it right, I think it, it presents this opportunity. You know, like at this point, I’ve integrated thing- tools, frontier model tools like Claude or, or ChatGPT into my daily work stream, right? Because it’s a super useful kind of personalized advisor on the types of things that I use it for.

 

[00:44:36] Michael Minkoff: And so ideally, and, and it has this li- you know, this global library of kind of “knowledge” that it can infer from, and then that gets complemented by the personalized exchanges that I have with it, right? And so if you get that at the farm level for, for these, you know, and, and it’s something we haven’t explicitly stated, I don’t think, in this conversation, but is generally well, well-known, right?

 

[00:44:57] Michael Minkoff: There’s a huge extension gap, [00:45:00] right? For that– for the smallholder and tenant farmers, right? You’re talking one to 6,000, one to 10,000 in some, some places. Um, one, one extension officer for 10,000 small scale producers or, or, or tenant farmers. Um, and so a lot of these farmers don’t necessarily have the advisory access, generally speaking.

 

[00:45:20] Michael Minkoff: And then if suddenly they can have a personalized assistant that can tell them, not just like you said earlier, David, how to plant, you know, rice more effectively, but actually why they should be focusing more on their livestock than, you know, farm crops this season because of market conditions, anticipated weather, and who knows, geopolitical realities and, and global ma- you know, whatever factors become integrated into this.

 

[00:45:44] Michael Minkoff: That becomes a super powerful and, and potentially very meaningful livelihood resource. And, and another thing we haven’t quite said in this is, right, agriculture operates in seasons, right? We’re not talking like a week, uh, or a month. We’re talking about six months, three [00:46:00] months, a year of someone living often in subsistence or near subsistence conditions being, you know, derailed.

 

[00:46:06] Michael Minkoff: That’s, that’s hugely… That’s a huge vulnerability, uh, and, and kind of moral hazard risk. So it– there’s real powerful potential and excitement there. Yeah. On the, on the worry side, um, building these systems is very hard, and build– especially in the context we’re talking about with the level of integration at kind of, you know, multi-state, multinational, mul- you know, global levels, uh, that gets sustained, uh, especially under kind of very dynamic geopolitical, uh, times and, and circumstances.

 

[00:46:40] Michael Minkoff: So, so I, you know, I worry about, um, if If the, and data remains fragmented and biased, then we start putting out digital tools, which agriculture has done many times in the past, and farmers are like, “This doesn’t actually help me.” And then it just becomes a lot of investment that doesn’t [00:47:00] lead to- Well, and you’ve 

 

[00:47:00] David Bergvinson: betrayed the trust of the farmer

 

[00:47:01] Michael Minkoff: and you’ve betrayed the trust of the farmers yet again with your newest, latest, and greatest digital attempt, uh, attempted digital intervention. So the, yeah. 

 

[00:47:09] David Bergvinson: Good, good points. And then Sowmya, you, final word, um, boots on the ground perspective. Love to hear your perspective on these two questions. 

 

[00:47:18] Soumya Alamaru: I totally agree with João and Mike.

 

[00:47:20] Soumya Alamaru: Uh, if we get the data corpus right, we can give truly personalized advisories. Not district level, not re- village level recommendations, but plot level advice that accounts for your soil type, your variety, your irrigation access, your historical pest patterns, and your mar- market access point. I want to give one example that we did recently, uh, uh, in Andhra Pradesh.

 

[00:47:48] Soumya Alamaru: We are trying to get plot level granule, uh, granular data. In, in one of the regions, which is a banana rich region, uh, most of the [00:48:00] farmers are, uh, har- uh, have banana plantations. We try to collect the data on which farmers are using fruit care activities, like bunch covering or injection, giving injections for, to get a, a fruit quality.

 

[00:48:18] Soumya Alamaru: To our surprise, we came to know that out of 600 farmers, there are only 12 farmers who are following these measures. Rest of the farmers don’t even know, or they’re not aware of, or, uh, they don’t know how to do it. Now, the, if, if we get this data that out of 600, 12 farmers are only following the fruit care activity, that granular data can truly, uh, target those rest of the farmers who are not following fruit care activities, and then we, the guidance can go to them how to improve the food, uh, fruit quality and gets, get better price in the market [00:49:00] So data cor-, the real data corpus like, uh, real, um, ground level data corpus can give such advisory.

 

[00:49:09] Soumya Alamaru: And second point on tenant farmers and women farmers are tribal farmers who have been excluded from land records can be visi- can m- may get visible through alternative data like input purchase records or market transaction history. So we can bring them into system by covering these data sources. Yeah.

 

[00:49:32] Soumya Alamaru: Third point I want to s- uh, talk from the government perspective, governments can have real-time granular data on what farmers are actually doing, not what they’re supposed to be doing. Policy can be adjusted mid-season rather than waiting for annual assessments, estimating losses and gains, and then taking an ad- uh, decision.

 

[00:49:56] Soumya Alamaru: What worries is if [00:50:00] the data, the data sys- models or AI systems are not built on, uh, representative data, AI will deliver high quality personalized advice to large connected progressive farmers only, and generate potentially harmful advice to smallholders. That widens productivity gap. That widens wealth gap.

 

[00:50:27] Soumya Alamaru: It will increase the inequalities which are already existing in the agriculture, uh, field. AI becomes another mechanism for agricultural polarization. 

 

[00:50:41] David Bergvinson: Yeah. We don’t wanna create a digital divide 2.0, so, you know- Correct … responsible collection use and delivery of data insights to farmers, policy makers, and value chain actors is critically important to unlock the value of this.

 

[00:50:54] David Bergvinson: Correct. Well, uh, thank you to all three of you for a very insightful conversation. Soumya, thank you for your boots [00:51:00] on the ground perspective from Andhra Pradesh. Make, from your broader perspective of translating this into projects, policies, and, and i- impact. And Xiaoyu for wrangling data across a wide range of international organizations and delivering value to national stakeholders through data models and, and wide range of tools.

 

[00:51:21] David Bergvinson: Thanks to all of you. You’ve raised some actually interesting questions for follow-up conversation, especially around personal identification information and other topics that we’ll follow up, uh, in Grounded Intelligence. So thank you all, and, uh, uh, very exciting journey that we’re on, and thank you for your contributions towards making it impactful and real for smallholder farmers around the world.

 

[00:51:49] Michael Minkoff: Today, 

 

[00:51:50] David Bergvinson: we’re gonna be talking about apples and not what you think. We’re gonna be talking about the diversity that exists in an orchard o- of apples, and [00:52:00] that no two apples in that orchard are exactly the same. That variation depends on the variety, depends on the microenvironment in which that apple was produced.

 

[00:52:11] David Bergvinson: It depends on the soil, the level of water, uh, how it was cultivated or pruned by the farmer. So when we talk about variation, not just in apples, but variation in our data systems to support modern agriculture, we’re talking about data that contextually is very different. And we ne- need to make sure that that variation is captured in the data that’s going into training models, especially models to support advisory services for farmers.

 

[00:52:39] David Bergvinson: So that means making sure that we capture the diversity of production systems, farmer profiles, especially women that are often underrepresented in these data assets for a variety of reasons, that we capture variation in the soil type in which these crops are growing. And also making sure that that [00:53:00] data that is being captured has associated with it what we call metadata, which is sort of smaller bits of data that actually describe the context of how that data was collected, um, the parameters that would help you understand how to best use that data in order to deliver targeted services to farmers.

 

[00:53:17] David Bergvinson: So when we talk about data systems, let’s talk about the diversity they represent, making sure that diversity captures the diversity of farmers and production systems that we want to deliver advisory services to, and to make sure that we’re getting the feedback so that w- outputs from those advisory services are constantly improved over time.

 

[00:53:38] David Bergvinson: And that in itself is a different and diverse data set. So that’s it for today. We’re gonna dive into this in more detail. See you on the next episode.

 

[00:53:54] David Bergvinson: So in our next episode, we’re going to be looking at moving from data, which was our [00:54:00] topic today, to models. And models are really what transforms that data into valuable insights that can inform decisions, can, uh, enable things to be done, uh, semi-autonomously or autonomously, autonomously on our behalf.

 

[00:54:15] David Bergvinson: So moving from data to models is a critical conversation, and we’re going to, um, be actually, uh, diving into detail on this topic through our discussion papers, which you can find on agxai.com. So join us for that following conversation on models. And until next time, take care.