[00:00:05] David Bergvinson: Well, welcome to episode four of Grounded Intelligence, and today we’re tackling a very challenging but important topic of benchmarking. Now, you might not know what that word means now, but at the end of this podcast, you’re gonna be almost an expert on this topic. And it really is focused on how we measure the performance of artificial intelligence in a range of applications, and our particular one of interest here is around advisory services.
[00:00:41] David Bergvinson: And I can think of no better guests to have on to talk to this topic than the Athena Info- Informatics team. The Athena Informatics team that is really leading the way in bringing together diverse partners to tackle this challenging [00:01:00] topic of benchmarking, and also leading a, a discussion paper on the same topic.
[00:01:04] David Bergvinson: So with me t- this morning is Praveen, uh, Pankajang. Uh, sorry, you’re– Uh, we’ll skip over these parts. And, uh, Deepa Karthikainen, uh, who are leading us– these efforts in, um, benchmarking, especially in India, where, uh, artificial intelligence to serve smallholder farmers is very advanced and this topic is very pressing.
[00:01:30] David Bergvinson: So, uh, with me, uh, is, uh, Praveen and Deepa, and so, uh, would you mind just both introducing yourselves a little, uh, so that our audience knows who is speaking to them on this topic?
[00:01:43] Deepa Karthykeyan: Praveen, would you like to go first?
[00:01:45] Praveen Pankajakshan: Um, sure. Yeah. Um, thank you so much, David, for hosting us. Really lovely to be part of this conversation and, um, yeah, and, and also talking about something which is I’m passionate about. So myself, uh, Praveen [00:02:00] Pankajakshan, and, uh, I’m a senior technical advisor, uh, here at Athena Economics. Um, I’m also a chief, uh, AI scientist, uh, at UrbanKisaan, apart from a few other, uh, roles and responsibilities that I have.
[00:02:13] Praveen Pankajakshan: Um, and, uh, yeah. And my background is, uh, um, machine learning and AI applied to a variety of problems and, uh, recently, um, for the past decade or so working on, uh, its application to agriculture. Yeah. Thank you so much, and over to you, Deepa
[00:02:32] Deepa Karthykeyan: Thanks, Praveen. Um, everything I’ve learnt, I’ve learnt on the job working with Praveen, so I think this is real privilege, Praveen, to also have you here and to be on this d- in this discussion. So David, also thank you, um, for including me. I’m probably gonna learn more than I share. But, um, Deepa Karthikeyan, I’m co-founder, partner of an organization called Athena Infonomics, and we’re, an [00:03:00] organization that works at the intersection of data tech and social impact. And increasingly, um, artificial intelligence is part of the intervention toolkit governments when thinking about, advancing social impact. So we’ve been going on a steep learning curve as an organization l- because so much of our work’s about either designing a program to make sure it delivers what it’s intended to or evaluating if it has.
[00:03:28] Deepa Karthykeyan: Um, so yeah, really that’s the experience that I hope to bring to this conversation.
[00:03:34] David Bergvinson: So that, that’s excellent because we’ve got deep technical expertise, uh, Praveen, that you bring to the, the table. And, uh, of course, Deepa, your, uh, team’s engagement with governments, private sector, and a range of organizations from health to farming systems to finance, uh, is, is wonderful because it’s, uh, that interface with true [00:04:00] application and the challenges we face in delivering these fast-moving, uh, technologies to a wide range of stakeholders.
[00:04:06] David Bergvinson: Um, so Praveen, maybe I’ll just turn to you, uh, as we paint the context of, uh, benchmarking in the context of advisory services for smallholder farmers, especially those in India. But it has a, you know, a, you know, a, a parallel, or you can transpose this to other parts of the world, especially Africa or, or Latin America.
[00:04:31] David Bergvinson: So, you know, from, from your perspective, maybe as we think about a farmer using an artificial intelligence-based, uh, advisory service,
[00:04:41] Praveen Pankajakshan: Right
[00:04:42] David Bergvinson: what do we need to think about as we think about model performance and the role that benchmarking plays in helping us assess the value, accuracy, relevance, um, that all come together to, to build trust and manage risk of these [00:05:00] advisory services to farmers?
[00:05:01] David Bergvinson: So, um, what goes through your mind as a technical expert in this, um, um, area of benchmarking?
[00:05:09] Praveen Pankajakshan: Right. I mean, it’s always, uh, interesting, right? Like, uh, and we have to look at how the, uh, whole benchmarking process itself evolved, right? Like, and traditionally it has been used for, uh, actually, like evaluating models itself, right? And so, um, even, uh, you know, when we actually like deploy it on the real world scenario, not only in India, but we have also done it in Africa, Latin America, and across the globe. Um, you know, and then, uh, one of the important elements is, uh, is actually how does it generalize, right? Uh, you actually train it for– obviously, like we don’t– when you look at models and almost everyone, right? Doesn’t have exposure to all the possible diverse data sets that you can actually find, uh, globally. Uh, so it’s n- it’s almost an impossible task to expose the [00:06:00] model to such a diverse, uh, data set. So, so the main question is, like, how do you actually like, uh, evaluate, uh, some of these models and frameworks for, uh, even data sets that which you have not, uh, looked at, right? So, um, so in terms of benchmarking, uh, you know, there are two things that comes to my mind is one is data set and the other is the evaluation metrics itself.
[00:06:23] Praveen Pankajakshan: So, uh, so both of these actually play an important role. Uh, and when we deploy like adv– maybe it could be some advisory services or it could even be some, uh, you know, computer vision models, uh, or even remote sensing or satellite imaging, uh, services. Uh, the thing that we are looking at is, uh, uh, can you excuse me.
[00:06:46] Praveen Pankajakshan: Can you actually expose the model, um, uh, during evaluation itself to different possible scenarios, right? Uh, different agro, uh, zones, uh, um, or, [00:07:00] uh, you know, weather conditions and everything. So, uh, those are some of the things that we look at, and it’s always a challenging, right? But how do you sample from this diversity and give this heterogeneity, uh, if I may say so, uh, to the model, uh, and expose that and hence, um, you have a s- like a kind of a, a closed environment of testbed, uh, that, is kind of a sub-sample of what you would actually see in the wild, right?
[00:07:28] Praveen Pankajakshan: So that’s actually mostly what I think about when I think about, um, uh, evaluation before wide scale deployment, you know. So, so I don’t know if I answered the question, but,
[00:07:39] David Bergvinson: Yeah.
[00:07:40] Praveen Pankajakshan: yeah.
[00:07:40] David Bergvinson: I,
[00:07:41] Praveen Pankajakshan: Yeah.
[00:07:41] David Bergvinson: and I, I’d like to, uh, you know, turn to Deepa who’s, you know, also interfacing then with a wide range of stakeholders to transpose what you’ve just described in your considerations of benchmarking. And, and Deepa, you know, how, how do you communicate this topic to, to your stakeholders in government or private sector?[00:08:00]
[00:08:00] Deepa Karthykeyan: I think it’s such a– So, I think there’s value in contextualizing where conversations around benchmarks sit in the context of safe and responsible AI use first, um, David, right? So one of the things we’ve observed, and again, a lot of what even in the work that we’re doing, there’s more time and money going into building models at the moment than these models actually being used So that feedback loop hasn’t necessarily kicked in, at least in the programming we’ve been engaged in so far. what’s already becoming very evident and what governments care about is what’s changing for the farmer, right? What is the impact? Have they been able to cut down input costs? Have they had… What’s changed? And I think there are two parts to it. is a lot of attention today [00:09:00] and the ways in which I think even in a context like India, the ecosystem is coming together to evaluate models is very much in sort of the pre-deployment or the controlled deployment phase, right? Which is you could have an independent academic institution that applies some of these standard evaluation metrics, and this goes back to what Praveen described as, um, so it’s important what metrics we want to consider. Um, now, governments would lead with, “Is model evaluated and certified?” And the answer could, to that could be, yes, it is, because there was a third-party academic institution that did certify it, but that was done on a pilot of 100 farmers in a pre-deployment or a deployment in the context of a controlled environment. Now, that doesn’t necessarily answer the question that I think governments truly care about, which is, what’s it changing and is it safe, right? And is it, um, relevant? [00:10:00] Uh, lot of that conversation, I don’t think benchmarks even start to address, right? Because you’re going into the domain of post-deployment monitoring for safety and risk. So when I think about governments, I think now for them, the focus is much more on the user side. It’s not necessarily on the, oh, this is a X-certified, uh, model because they had the perfect blue scores or F1 scores or whatever it is that we’d want to use to say that. Uh, that’s, in my view, not necessarily what governments use as their gold standard for evaluating what is and isn’t working. So there is something to be said about where benchmarks even fit in the real world and in the context of what governments truly care about.
[00:10:48] David Bergvinson: Well, this is teasing out just the very issues that I hoped for in this conversation, so thank you. Um, you know, and, and I think you’re right. Governments wanna have the assurance that it’s been tested, [00:11:00] but the true test is in deployment. And actually, maybe even we need to think, rethink benchmarking in post-deployment and how these models evol- evolve over time to better serve end users.
[00:11:12] David Bergvinson: So, uh, a lot of really good points brought up, uh, by both of you already. You know, one of the challenges with developing large language models is that most of them are coming from advanced economies or from the Global North, however you wanna label it, in which, uh, there’s a biased sampling of that data.
[00:11:28] David Bergvinson: It tends to be from organizations based in Europe or North America. You know, it’s drawing on commercial agriculture for large scale farms. Uh, and, uh, you know, these farms, uh, tend to grow specific crops like, you know, uh, maize or soybean. And now we’re pushing them into serving farmers that are on small scale enterprises, you know, less than a hectare or, or two acres.
[00:11:58] David Bergvinson: They grow a wide range of [00:12:00] crops. Uh, they have livestock as part of the mix. They have alternate sources of income. You know, it’s a very different context. And so, you know, these models, you know, need to integrate other data assets, which is another topic, uh, for another day. But the end result is that we are, you know, using, um, benchmarking tools that have been developed in the Global North to measure these models that are being applied in the Global South.
[00:12:30] David Bergvinson: And so, you know, this presents a real challenge. Praveen, maybe you could speak to that issue of, you know, the data going into the development of these models. How does that impact on the testing or benchmarking of these LLMs, and are there specific new tools we need to benchmark in the context of India, for example?
[00:12:50] Praveen Pankajakshan: Right. Absolutely. and you know, even piggybacking on, um, what was communicated or what was discussed earlier, um, is [00:13:00] that, uh, in fact, like, um, like you rightly pointed out, um, even, uh, even from a remote sensing or satellite imaging perspective to LLMs, uh, to computer vision, uh, many of the dataset that is chosen is, uh, uh, obviously like, um, is very focused, right?
[00:13:19] Praveen Pankajakshan: Like, and, um, so when you train the model on that, like, um, it’s– and when you take it to a small landholding size, uh, the challenge is very different. Like, just to give you an example, um, when we were testing some models, you find that within a particular land there is not only the farm, uh, but you also have a small cottage, or you have a small, uh, road passing through and like, um. So when you are looking at land parcels, then you’re looking at land parcels as a whole, which include these components as well, right? Like even a livestock, an area for a livestock. So, uh, when you’re doing, uh, some kind of crop [00:14:00] identification and things that, uh, it becomes really challenging, uh, to do it in those parcels.
[00:14:05] Praveen Pankajakshan: So, th-uh, the models themselves, they assume that the parcel which is given is actually having only crops, right? So then you need to probably like retrain the model, uh, or even build the framework itself, right? From ground up. Um, saying that, you know, you are going to ex- uh, expose the model during, uh, inferencing and deployment to, uh, not only, uh, croplands, but you’re going to have like other, like there’s going to be a bore well, there’s going to be other things, right?
[00:14:36] Praveen Pankajakshan: So in some way, I think there needs to be a thought process also, uh, when you build the model. Now I’ve come to an important point, is the design framework itself. So the, uh, evaluation hence has been thought of only from a certain perspective, right? Assuming that, uh, for example, the land behaves in a certain way, and so the evaluations are done in that.
[00:14:58] Praveen Pankajakshan: But, uh, if you just step back [00:15:00] and look at what Deepa was saying or what you were saying, like we have not actually designed systems or frameworks or evaluation systems looking at from a perspective of what is the impact measure that we wanna have. We have looked at F1 scores but even, uh, most of the recent evaluations that we are looking at bias, fairness has been thought of from an as an afterthought, right? Uh, and that is because of how these models has been traditionally been evaluated as we have been looking at it from a very unidimensional problem of like either, uh, error rates or F1 scores or BLEU scores and things like that, right? So, um, I think that definitely needs a of like, uh, now that we have like large scale, uh, systems deployed, now we have to go back to the, uh, whiteboard and, uh- You know, like design new, uh, frameworks and even evaluation metrics, I would say, which combines all the elements that you talked about, uh, David, at the beginning, um, [00:16:00] including bias, fairness, um, right, and all those elements.
[00:16:03] Praveen Pankajakshan: So how do we, uh– And when we are doing large scale deployment is when we get these gaps, right? Like we find these gaps, and we know that benchmark datasets have these gaps, and hence evaluations have it, models have it, right? So how do we now… Because after now that we have like wide scale deployment, we are seeing those gaps. How do we now go back and reframe those, right? And redesign. So that means like the whole thought process of right from design stage to building the model to evaluation has to be rethought about.
[00:16:36] Deepa Karthykeyan: And
[00:16:36] David Bergvinson: Yeah. Yeah, and, and, you know, Deepa, your thoughts on this because, I mean, you’re, you’re interfacing with end users, a, a wide range of them. Um, and, you know, they’re saying, well, you know, how relevant and accurate and what are the risks or liabilities of using these things, you know, especially from a government perspective?
[00:16:56] David Bergvinson: Um, and, and does that feedback go, [00:17:00] you know, back to people like, uh, uh, uh, Praveen to, you know, refine the approach or methodology of benchmarking?
[00:17:09] Deepa Karthykeyan: So that’s– this is such an interesting question because it leads with the assumption that there is a Praveen that’s mandated and paid and resourced to be thinking about these questions of customization. And I think that there is an incentive gap there, right? Because today, and again, and you know, there was a different conversation where we were talking about what does scalability mean in this, in the context of LLMs and, applications that are built on it.
[00:17:37] Deepa Karthykeyan: And of, in the traditional tech world, it was taking one product that could, with very minimal customization, serve a very large number of customers. AI changes all of those rules in such fundamental ways, David, right? It’s adaptive. It engages with different users differently. And in this, a sector like agriculture, which introduces a wholly new set of complexities because [00:18:00] agronomy so much, right?
[00:18:03] Deepa Karthykeyan: And I serve on the board of a crop insurance company, and we– I mean, there are so many different microclimatic gates that we serve and insure, and even just within Andhra Pradesh, it is so complex and diverse. So the number– I mean, AI in some ways is very powerful in solving for it. But then the question for me that I think, and I’m taking one step, I’m, I’m, I’m zooming out a little bit and asking the question is there enough private sector incentive to build applications at that level of customization?
[00:18:34] Deepa Karthykeyan: And who’s doing, who’s doing that? Or should governments come in and underwrite that risk of leaving certain groups or certain sets of users behind because there just isn’t the right set of upstream incentives train at that level of localization? I think I start with that question of who is resourced to be thinking about, uh, some of [00:19:00] the things that we’re just describing here and talking about in terms of benchmarks and that level of customization and relevance.
[00:19:06] David Bergvinson: Well, well, Praveen, is there– I, I mean, I’ve been exploring a lot on skills, you know, using skills to populate, uh, different commercial, uh, AI services to enhance the output performance. Is there a thought around building skill profiles for benchmarking tools that then address some of the points that, uh, Deep has just brought out around hyper-specialization for niche end uses?
[00:19:36] David Bergvinson: Um, because, you know, as a human, it’s hard to do all that, but then how do we leverage AI to actually improve benchmarking to serve these diverse audiences and applications?
[00:19:46] Praveen Pankajakshan: Right. And I think there is a lot of conversations happening around that, but it’s, it’s a very, um… I would say it’s, there is a lot of nuances to that also, D- David, because, uh, uh, you know, human beings [00:20:00] themselves are so diverse. Like, um, we were looking at, for example, evaluating, um, some preferences towards the output of certain models and, uh, even, uh, with respect to outputs which are provided already by a human, uh, evaluator and evaluating the human evaluator by multiple evaluators themselves, there is not an agreement Uh, because like, uh, as human beings, we have preferences over one thing over the other. So usually what we do is we take, um, um, either if you’re looking at, um, a data set which has been created by, uh, uh, by people or by, uh, created, um, you know, there are some generated data set by, um, uh, by AI models as well. Uh, when we are evaluating that, we just don’t look at whether it, it may be human evaluators or AI model evaluators. We just don’t look at, uh, uh, just having, like, one particular evaluator, right? How do– how does it [00:21:00] perform? How– What are the kind of preferences that are there over multiple evaluators? then you kind of, uh, ensemble them, bring them all together to get a clearer picture, right? So statistically, when you do that frequently, uh, as much frequently as possible, then you, uh, the picture, uh, you know, becomes more and more clear. Um, AI models have their own, like, the only challenge is using AI model as an evaluator, um, and like aggregating them over a period of time, is that they do have their own, uh, biases, right? And so, uh, inherently, they all seem to agree on certain points and disagree on certain other elements. but it– you are kind of like in a way, um, uh, f-forcing to, uh, reinforce a certain kind of, uh, uh, bias and belief which is already, uh, hidden and, uh, so there is high risk that that, um, doesn’t actually get picked up early on.
[00:21:54] Praveen Pankajakshan: So it’s always good idea to have like, uh, random selection and having that, [00:22:00] um, evaluated with multiple human evaluators. But then the major question is that, how much of, you know, time and effort is required, uh, by, you know, to do that, right, at a human level. So it’s when you’re talking about like millions of queries that are coming in, it’s always difficult, right?
[00:22:16] Praveen Pankajakshan: So, so, uh, obviously after some time, it’s going to reach a certain stage of, uh, exponential saturation, right? Um, but even to your question of, uh, is that, are there some new skills that are emerging? Definitely. Like, um, I think this whole idea of looking at this entire design framework, uh, is very important. Um, and then, uh, what I, you know, as part of this, uh, whole topic itself, I was thinking like this is going back to the table and thinking about like even what do we have as evaluations, right? Like I come back to the metric because it’s an important, important tool. Um, and the metric system itself, we have not had like robust metric systems for evaluating something like, uh, you know, a [00:23:00] multiple of factors like, uh, you and Deepa were pointing about to farmer income or farmer impact, um, you know, even environmental impact trust, right?
[00:23:09] Praveen Pankajakshan: And so there are a lot of areas that really need to be addressed, right? And how do we address them? Um, reliably. And that involves like a multiple set of people coming together, not only data scientists or AI scientists, but it also needs field teams coming in, policymakers, economists, and all of that.
[00:23:28] Praveen Pankajakshan: Environmentalists, right? Ecologists coming together. So, um, agronomists. So all of these inputs need to be taken when we want to design a robust system, right? Like, so, I think that’s what is of major importance
[00:23:42] David Bergvinson: there a gold standard currently for benchmarking in, in the ag domain?
[00:23:47] Praveen Pankajakshan: Um, there are with respect to, um, I think like, um, certain datasets, uh, it’s been in silos built by different entities. Uh, so there is a concerted effort, um, in the open source, at least community, [00:24:00] to bring all of these together. Uh, so there are some common initiatives, uh, which are, uh, very significant. There are some, um, I think like benchmark when you talk about some foundation models.
[00:24:11] Praveen Pankajakshan: I think, uh, there are– have been released, some released in the, in the domain. But, uh, where we want to get to, I think that raises an important question as well, uh, David, is that most of this is released by private entities, right? Like, and think there needs to be also a global concerted effort to get these, uh, benchmarks, datasets, and governance, all of them put together.
[00:24:34] Praveen Pankajakshan: We’re not talking about AI governance, right? But, um, even like, uh, we have to talk a little bit about impact governance as well, and even usage governance and dataset, uh, sovereign datasets and things like that, right? So these are, I think, important conversations that we have to bring to the table, and these are not yet there.
[00:24:52] Praveen Pankajakshan: These are, these are still like, you know, um, ha-ha-happening behind closed doors, but we need to bring that to the forefront. [00:25:00] So
[00:25:01] David Bergvinson: Deepa, I mean, from your experience of working with governments, do you see any, uh, you know, organization or unit that is really driving this standardization, be– you know, gold standard certification or aggregating these benchmarking tools in the context of India,
[00:25:20] Deepa Karthykeyan: Yeah. I
[00:25:21] David Bergvinson: let’s say?
[00:25:21] Deepa Karthykeyan: that we’re starting to see an interesting trend where we now have departments of AI state governments are starting to set up with a minister overseeing it, which mandate, the forms, and the functions that these departments are likely to perform is yet to be seen. But for me, the fact that it’s dem– governments are starting to take this s-seriously and to the point that it merits a dedicated institutional, m-ministry almost, is, I think, encouraging, David. Uh, I do see that the… And the current, at [00:26:00] least the current protocols, is that we are leaning on academic institutions that are certified, like AI centers of excellence, to be able to provide some of that guidance in terms of certifications. but I, I think we’re at that stage in the curve where as there is greater use and scale that becomes more visible and some risks start surfacing, right?
[00:26:25] Deepa Karthykeyan: Risks, both political as well as safety risks, become more visible. Um, and for example, they’re already very visible in health, right? So you’d see therefore health gets a lot more attention. It’s a lot more mature. Um, I would say followed by education and then perhaps agriculture in that order in some sense, right?
[00:26:44] Deepa Karthykeyan: Uh, the, the risks with ag or agriculture are still less visible. Um, but I think when that changes, uh, I’m expecting to see that governments will start taking on a regulatory role, [00:27:00] um, and will take on that role a lot more seriously, and perhaps also build capacity within government institutions
[00:27:06] David Bergvinson: Are, are there lessons we can learn from the health, uh, given it’s more advanced and apply it, uh, to the ag domain?
[00:27:15] Deepa Karthykeyan: I think there’s de– I– there’s definitely value in looking at the institutional checks and balances that the health sector brings, and I think a lot of how– again, health is different, right? Like a drug interaction is a drug interaction, right? Uh, it’s, it’s– agriculture can be more complicated, particularly on the advisory side in ways that health may not be.
[00:27:36] Deepa Karthykeyan: I’m not a health expert, so I wanna caveat that with that sort of I’m, I’m, I’m, guessing, uh, uh, health is less complicated, at least in terms of causality and thinking through some of those bits than agriculture is. health also has been more intensely regulated historically, right? They’ve had stronger regulatory authorities. Um, and that sort of institutional checks [00:28:00] and balances, I think, has created a space where it’s been easier to come in and think about benchmarks and model evaluation and safety with a lot more seriousness and institutional capacity than I would say a sector like agriculture has. So maybe there’s something there in terms of looking at the institutional checks and balances and models and the roles and responsibilities of regulators, for example, and what they’ve done and what policies and standards and processes or SOPs they’ve put in place.
[00:28:30] Deepa Karthykeyan: I think there’s something agriculture can certainly learn from that.
[00:28:33] David Bergvinson: You’re, you’re– I, I agree. You know, agriculture is very complex. It’s very context specific. You know, it’s extremely challenging. It’s, it, it’s a data intense enterprise, I say, but it’s operating in a data poor environment. And so I think that translates to AI as well
[00:28:49] Deepa Karthykeyan: And the one thing I do, because I also want to be optimistic, and one of the reasons I’m excited about ag, David, while it is complex, and this is maybe a, a call to entrepreneurs and [00:29:00] enterprises, the farmer is an entrepreneur, right? He’s a micro– or she’s a micro entrepreneur. They’re optimizing costs.
[00:29:06] Deepa Karthykeyan: They’re trying to increase revenue. They’re trying to do it often in, in an environment where there are serious external risks that they don’t fully control. So there’s– they’re probably– So– And they understand and are directly incentivized by wanting to create economic value. So there is a real there, right?
[00:29:26] Deepa Karthykeyan: That can be capitalized in a market-based response for entrepreneurs to come and be responsive to our enterprises and, and then governments to underwrite safety and responsibility on. So there are things about agriculture that I think are also very unique and powerful in ways that health may not be, where the beneficiary, um, angle of framing the user is maybe more dominant, while in agriculture it’s sort of really, you’re really a customer in a lot of ways because there’s a direct economic, uh, gain, uh, that I
[00:29:56] David Bergvinson: Well, thi–
[00:29:57] Deepa Karthykeyan: think is more attributable.
[00:29:58] David Bergvinson: well, this conversation has sparked [00:30:00] some thinking in me. I, you know, I thought I came into this conversation knowing enough about benchmarking, but I realize there’s a lot more we need to discover here and, um, and, and to be humble enough to, uh, you know, learn as we go, while at the same time recognizing we need to be responsible in using best-in-class benchmarking tools to at least say, “Yes, it has been tested,” but recognizing that when it goes into the real world, that’s where the real testing takes place, and putting in place the instruments to allow us to vet and be confident it’s delivering the intended goal of improving income, reducing risk, increasing equity and access to knowledge, et cetera.
[00:30:42] David Bergvinson: Um, from our conversation, just to wrap up, I, uh, I’ve got my own list of things that I need to follow up on. What did you take away from this conversation, Praveen, that you would consider doing differently?
[00:30:54] Praveen Pankajakshan: Um, well, I mean, it’s, uh, it’s an ongoing journey, d- uh, David. Um, I think there’s [00:31:00] lot, uh, lot, lots to learn and lots to catch up on. I, I mean, we talked about standards, right? I think it’s, it’s very interesting, um, you know, like even talking about, uh, uh, data sets and, uh, data standards and interoperability like, um…
[00:31:17] Praveen Pankajakshan: So I think like, uh, definitely there is, uh, uh, uh, there is a lot of thought process that is required in reflecting on that. And, uh, I would also maybe go back and… I mean, one of the things that I’m really excited about is even like, uh, building something from first principles, like, you know, I’m like starting to think like, what is that metric that’s a universal metric that goes beyond the F1 scores that one can probably invent, like, uh, you know, which kinds of combines together like, uh, trust, uh, reliability and, uh, uh, even, uh, scaling like with different geographies and equity and all these, uh, other elements that you talked about, right?
[00:31:59] Praveen Pankajakshan: So, [00:32:00] uh, is there some way to kind of even standardize that, right? Like, uh, because, uh, trust is a very– uh, some of these are very gray areas, right? How do you still, uh, figure out mechanisms for quantifying that, right? A qualitative measure, but quantifying that in a reliable way and making that into a global standard, a gold standard.
[00:32:22] David Bergvinson: Mm-hmm.
[00:32:23] Praveen Pankajakshan: I think that’s something to think about and reflect on. So maybe I’ll,
[00:32:26] David Bergvinson: That, that’s a good challenge to leave us with.
[00:32:29] Praveen Pankajakshan: Yeah.
[00:32:29] David Bergvinson: Deepa, what’s, what’s your insight, uh, from, from this conversation?
[00:32:34] Deepa Karthykeyan: Yeah. We’re, we’re sailing the boat as we’re building it, David. So I think it’s, it therefore it brings with it all of the complexity that that journey does. And I, think I’m… But I am encouraged by, by what I see in countries like India, where governments are stepping in and taking this more seriously. Um, I’m also encouraged by of the [00:33:00] recent, more sort of– Because the geo– I was always worried that the geopolitics will prioritize investing on the model side and not necessarily the resources needed for questions like safety and risk.
[00:33:17] Deepa Karthykeyan: But, and I see that changing too, uh, globally. So I think all of these are encouraging. Uh, we’re all gonna be learning along the way, and it’s a little bit of that learning by doing. So I echo and, uh, what Praveen described deeply resonates
[00:33:31] David Bergvinson: Do, do we all, do we sort of need a structure to help navigate or, and systematically approach this or recognize that, yes, we’re flying the plane while we’re building it, but does, is there a structure that needs to be in place to get to an endpoint?
[00:33:45] Deepa Karthykeyan: I do think that, I mean, it’s, it’s so difficult and even when I think about a country like India, every state has such a different mix of political economy incentives, institutional architectures, maturity, fiscal resources that it’s [00:34:00] almost… I- great question actually. I don’t think I, I wanna respond to that with like an answer because I think it’s a question worth about, saying, is there some sort of a method to this, this crazy navigating this?
[00:34:14] Deepa Karthykeyan: But I, I certainly think it’s… Yeah, I think it’s an interesting one to think about, David.
[00:34:18] David Bergvinson: Well, maybe that, that’s a good way to end with another conversation to follow. Uh, Deepa, Praveen, thank you so much for joining us on this, uh, topic, and it, it is a evolving topic, so I look forward to circling back and learning from both of you, uh, as we, um, apply benchmarking in practical terms to serve a wide audience in India, Africa, Latin America to empower farmers.
[00:34:44] David Bergvinson: But also, what came up in our conversation today was trust, and it’s one word that’s come up repeatedly in other topics. And so I, I see that as a thread, that benchmarking is a way for us to be a little more confident that we’re preserving the trust of end [00:35:00] users as we apply these tools. So thank you both, uh, and I look forward to following up with a subsequent conversation
[00:35:07] Deepa Karthykeyan: Thank you, David. Thank you so much. This is great
[00:35:09] Praveen Pankajakshan: Yeah.
[00:35:10] David Bergvinson: Keep up the great work. Thank you so much.
[00:35:12] Praveen Pankajakshan: you
[00:35:12] David Bergvinson: Bye for now