Go back

Accountancy Insights: LLMs and spreadsheets, Companies House IDV, and the OBR’s role

34m 11s

Accountancy Insights: LLMs and spreadsheets, Companies House IDV, and the OBR’s role

The transcription covers Accountancy Insights discussing the value and limitations of large language models (LLMs) for spreadsheet tasks. The conversation delves into the research conducted by Simon Thorne on LLMs and their performance in various spreadsheet tasks. It highlights the inconsistency in LLMs' abilities and their limitations in more complex scenarios. The discussion also touches on the Companies House ID verification requirements for directors and persons of significant control, emphasizing the need for compliance and the distinction between AML supervision and ACSP verification. The conversation reveals potential challenges with overseas directors and the importance of thorough ID verification. Overall, the transcription provides insights into the evolving landscape of technology in accounting and the regulatory requirements facing professionals in the field.

Transcription

5470 Words, 31287 Characters

Hello, welcome to Accountancy Insights. Three topics today. And first up, AI will be looking at the best and the worst large language models for spreadsheets, and finding out if any of them are truly dependable. Next, with Company's House ID verification just around the corner, what are the key need-to-nodes for accountants and to accompany directors? And our third item? Well, that is something of an explainer. We all know how important the Office of Budget Responsibility is. Its economic forecasts are always in the news, and doubtless cost the Chancellor a good night's sleep from time to time. But how much do you know about the OBR's precise function and how it actually works? With the autumn statement next month, it feels like the perfect time to find out. And just before we start, a reminder that the time you're about to spend listening to this podcast, it counts towards your annual CPD. So log your listens on the ICAW website. It is very quick. And subscribe to the podcast on your preferred app so that you can take advantage of every single episode for your CPD record. Let's talk about spreadsheets and AI. Simon Thorne, Senior Lecturer in Computer Science for Cardiff Metropolitan University, is joining us from Cardiff to talk us through his research he's been doing on this. Hi, Simon. Hi there. Good morning. So tell me about this study you've been doing. What were you hoping to find out? I was trying to understand what the real value of LLMs are in spreadsheets. This is a topic that I've been interested in since ChatGPT first came about and was a major release. What I was interested to know was how good is this at doing real world tasks. There are plenty of benchmarks out there for spreadsheets, but they tend to focus on very narrow, prompt formula type structure. And I was interested to know beyond that, could it do more complex work, more realistic kind of work. And that was the idea behind this. It was also to consolidate work that I'd done in 2023 for USPRIG, which is the European Spreadsheets Risks Interest Group, where I'd presented an initial examination of ChatGPT for different spreadsheet tasks and found that its performance wasn't very consistent. And in fact, it failed many of the tasks. So I suppose, you know, a lot has happened since then. And this was an attempt to try and encapsulate that and understand what the real value might be to people who want to leverage this technology to do their work. You looked at it as you say a couple of years ago. This is a field that moves fast. What did you expect to find this time? Did you think things would be much better? Certainly, I thought that things would have moved forward. I could tell from interacting with ChatGPT and other LLMs that their general ability to respond to queries and more complex queries had definitely improved. You could also do things that you couldn't such as link a document directly to it. So obviously spreadsheets can be loaded directly into it. So I was interested to know what capability it had to perform a more holistic analysis on that document and were there any limitations to how that is technically implemented. Obviously, we can't go into all your findings right now, but do you want to just run us through all the major issues that you're still seeing? In essence, the issues are really still the same from my earlier paper. They're better at answering in a more holistic way, but I tested a number of different approaches. So for instance, there's a task that's in the academic literature called the wall task, and this focuses on a very, very simple spreadsheet creation type task. When I tried this in 2023 with the initial versions of ChatGPT, it was utterly unable to answer in a coherent way. When I tried it again in 2025, I found that almost all of the LLMs, so I had a whole panel of different LLMs, could answer that correctly. So that's encouraging? Yes, absolutely. However, when I started to introduce more novel work to it, it was less able to respond. And broadly speaking, across the range of different LLMs, the abilities of these models is quite fractured. So some are good at doing certain things like text handling. Others are better at creating formulas. Others are better at, say, checking a spreadsheet for errors. But even within that, there's quite a bit of inconsistency. So I had a spreadsheet, a very simple profit and loss type spreadsheet with some mistakes that I'd intentionally put into it. You know, for instance, I'd omitted the costs of goods sold in the profit calculation in one part. I'd put a data entry error into the fixed costs of another. So quite realistic errors that might happen? Yes, that's what I wanted to go for, yeah. And it was patchy at finding those things. So it may spot a mistake in one part of the spreadsheet, but it wouldn't spot exactly the same mistake elsewhere. So it's inconsistent is what I found. And Why is that? Well, I think that inconsistency comes from perhaps limitations on how it reads the file that you give it. It may not read the entire file. That's the conclusion I came to, that potentially it has a certain hard limit that isn't obvious to the user as to how much of that file it can actually consume. And it gets to a limit and it stops, but it doesn't necessarily tell you that. And is there also an issue around the way these models are trained? And since they're kind of being taught to the test, is that right? That's correct. Yeah. So the LLM benchmarks are out there as a measure of success for LLMs. However, what's happening now is that these companies are training to these benchmarks. So it's slightly unsurprising that they do well in these tasks and they become, you know, it roads their value. I think we're seeing the same thing with the war task in as much that I ran it back in 2023. I think it's highly likely that that will have been picked up and incorporated into the training. And hence, when it comes to it in 2025, it has a much better way of answering novelty is a problem for it. So, you know, when I present something novel, so there's an extension of the wall task called the wall and ball, which is a much more complex calculation. It's about filling up a hot air balloon with helium and calculating surface area and complex calculations like that. So when I gave it that test, it didn't perform very well at all. And I assume that is because it hasn't seen anything like this before. This is how LLMs work. So back in 2023, which were the best and the worst then in 2023? I only tested chat GPT 3.5 because it was the only one available. So now then the kind of the most important question of all, which are the best and the worst now? Interestingly, the best model was Gemini 2.5 Pro as it was called at the time. Now I say surprisingly because Gemini had been the worst performer around the time that I did the testing. In second place was chat GPT at the time for Oh, all of that's changed now, of course. So chat GPT have, you know, released a new model and you can't reach chat GPT for anymore. Yeah, there's endless iterations. It's confusing, isn't it? It is, it is. So I mean, I guess suppose the other thing to say is that, you know, things change very quickly from release to release. And that was part of the thinking behind this benchmark was to have this as a modular repeatable test that we could use to come up with some objective judgment on, you know, which is best for this kind of work. But co-pilot, which is obviously the one, you know, it's the one a lot of people use, chat perform really poorly across all the tasks. Is that right? Yes, on the whole, it did very poorly. And to be honest, this is my experience of it in general. I think what's happened there is that it hasn't changed very much since the first release of it. And everything else is so they haven't invested in it effectively. I believe that's right. Yeah, I don't think it's been redeveloped. And if you look at all the other providers, they've released multiple versions of their LLMs. And there's real noticeable improvement in that. Co-pilot has improved a little bit in that period. But to me, it feels like it's still that kind of early model and is less capable than the others. So as you say, Gemini is out in front right now. Obviously, they had their difficulties last year and presumably they've invested heavily in improving it. I suppose it's a question of whether they're going to maintain that lead, isn't it? Yes. I mean, that will be interesting to see. I understand that Google plans to release the new model of Gemini, if it's going to be called Gemini, next week. So there'll be three points something out next week. And again, it'll be very interesting to see how that shapes up. Actually, by the time we release this podcast, that might be this week. So imminently? Imminently, yes, indeed. Absolutely. But you've not seen that? I've not seen the advanced version. No, I haven't seen that yet. But 2.5 is very good. It's perhaps the best. So that would be your recommendation right now? Yes, right now. That's what I would recommend. But what are the lessons here for accountants? Because none of them are going to stand out excellent across the board. The things you have to keep in mind when you're using these things. Number one, they can definitely save you some time. But you need to reinvest some of that time to validate and verify what the output is and to make sure that it hasn't missed anything. And that all the statements coming out of it are correct. So I would never trust these things blindly. You know, they can be useful assistance and they can do some of the work for you. But I would never assume that it's going to be fully comprehensive or fully accurate. And there's some issues emerging around the way these models handle your data? Personally, identifiable information should never ever be put into any LLM. The risk is that that information will be trained on and that information will come out somehow in a later release of a newer LLM that's based on that data. I mean, this is going to surprise people because I've seen the settings, there's often the opportunity to say, no, you don't want that to happen. But are you suggesting we can't really rely on that? Well, I think it's unclear exactly how that works. I would personally never do that, even if it says we won't, you know, we won't record your information. I think there's a chance, you know, if it's out there, it could potentially be used. And the phrasing in the terms and conditions is quite vague, as I understand it. I believe when you dig down into it, it sort of says that it won't usually be used for training. And if you go into ChatGPT as well, there's a way deep in the menus. There's a tick box, which you can tick, which says don't train from my data. But I personally don't, you know, trust that a great deal. And there's also a sort of safe mode as well that you can use, which says forget this conversation. But again, if you look at the details, it says, oh, yeah, it's deleted. But we might keep it for 24 hours. So it's, it's a little unclear exactly. LLMs are very hungry for more written discourse, you know, because that's how they work. And the size of these models mean that there's almost not enough human language to train on. So anything is valuable, essentially. These things, they save a lot of time. Obviously, you know, everyone's going to be using them. What's the best advice right now? I think the best advice right now is definitely use them. But you must check. You must validate and you must verify. If you do those things, then you can be confident in the output. And you can reap the benefits of the time it can save you. But be careful with it as well. So I feel like it kind of shifts us from being the primary workers to more the supervisors of these LLMs. And we must be sure that it's right because we wouldn't want to make those kinds of mistakes. Thanks, Simon. We will link to your study in the show notes, but thanks very much for being with us. Thank you very much. Onto Companies House ID verification. Now we've talked about it on the podcast before, but it arrives on November 18th. So ICAW's in-house expert, Mike Miller, is going to remind us what you need to know. Hello, Mike. Hello. Thank you for having me. It's imminent. Who does it apply to in the very first instance? Initially, it applies to directors of companies, members of LLPs, and persons of significant control of companies, so PSCs. There are some plans to broaden it somewhere down the line in terms of corporate members of LLPs, company secretaries, et cetera, in order to basically encompass everyone who has some sort of influence over the running of a company. But we're starting with the obvious people, the directors, the PSCs, and the members of LLPs from the 18th of November. So how does this ID verification work in practice? What do directors actually need to do? So for directors who are already with Companies House registered with Companies House, there's essentially a one-year grace period, implementation period from the 18th of November. But for anybody who's wanting to establish a company and register with Companies House, they will have to do it immediately, essentially, to be able to fulfill their duties as being part of the register. So there are a couple of ways of doing it, the first and probably easiest way, if you don't have an established relationship with either an accountant or a solicitor, is to essentially register on the government's own website to do your ID verification, provide a primary source of ID, which is generally going to be a passport or a driving license. If you don't have a passport or a driving license, we know that's a bit of a concern for older people, particularly if they don't travel or if they don't drive. There are other ways of doing it through secondary identification measures, such as your birth certificate, utility bills, et cetera. But essentially, you need to be complying with this, if you're going to establish a company, you need to do it as soon as possible. Thinking about accounting firms, is this mostly about them needing to remind their clients to get verified? Yes, I think so. If accountancy firms have clients who are directors or PSCs of companies, then they should definitely be at least knocking on the door and saying, "Look, you need to do this." Of course, if their existing clients are probably already established within company's house, so they will have this one-year grace period. But if anybody contacts their accountant and says, "I'm thinking of establishing a company," they will have to do it imminently, because essentially, you won't be able to register a company on company's house unless you complete your verification checks from the 18th of November. What do they need, then, in order to file on behalf of clients? At the moment, it's completely fine. You can continue filing as you would do for your clients. In the future, as it is told by a company's house at the moment, from spring 2026 for companies to file on behalf of their clients, they will also need to be registered as an ACSB and Authorized Corporate Service Provider, which means that they have to go through the verification checks themselves. This leads on to a few other technical challenges in terms of doing verification, whether you verify for your clients or whether you're just filing their accounts. But from 2026, it is expected that anybody who files accounts should be registered as an ACSB. And they're going to need director's codes, too? They probably will need director's codes. They will have to fulfill the obligations under the director's codes. The whole idea of this is essentially to get a unique identifier for anyone who registers on company's house. So this is every director, Mike? At the moment, it's every director who would be determined to be in control of a company. It doesn't cover, for example, large companies and those who have directors in their title, although that is something that is being explored by company's house and may come in further down the line. This is all done by secondary legislation. So statutory instruments that have been laid in parliament, it takes a while for these things to go through. And of course, the level of responsibility of a director in quotes of a large firm can vary very much depending on essentially what their responsibilities are and what they're doing. So I think that's a bigger challenge that is going to be explored by company's house down the line at the minute. They're really trying to, because this is such a large change for company's house and it requires a huge amount of resource and it requires a huge amount of attention to essentially plug the gap that has existed for quite a long time in terms of being able to register these companies. So it's kind of a bit by bit process to get it to a level of compliance that's desired by government. So thinking about ID verification right now, am I right in thinking there's some confusion between AML, supervision and ACSP verification? Yes. What's going on there? That is an issue we obviously, ICAW currently, is a supervisor for AML and CTF, Canada Terrorism Financing, which comes under the Anti-Modern Laundering Regulations. Now, company's house has determined that anyone who wants to register any firm, any person in the accountant who wants to register as an Authorised Corporate Service Provider has to be supervised by someone under the AML regulations, which I'm assuming is done by them to get a sort of baseline of compliance, but the difference is quite stark between what is required under the AML regulations, which is essentially a risk-based approach. So you don't need to check everybody's documents thoroughly. You do it on what you assess as the level of risk for that particular person that you're supervising. And ID verification, where you do need to legally go through everything and make sure that they comply, you know, it's a genuine driver's license or it's a genuine passport. And if you don't have the expertise, then you need to either use an automated system or you need to have some training in the ability to determine that it is. Which also raises some complications about, for example, overseas directors. And we've had this since their register of overseas entities was established a couple of years ago, as it can be very difficult to verify overseas documents if you don't speak the language, if you're not familiar with what the traditional forms of ID are from other jurisdictions, for example. So I guess the underlying message is just because you comply with the AML regulations doesn't mean you're doing verification to the necessary legal standard. And it goes out saying there is comprehensive guidance about all this on the ICW website. There is, we will be doing more and more both from my side, from our professional standards department have put out quite a lot of this guidance, professional standards to oversee the AML supervision for now, although that is potentially changing because of the government announcement yesterday. So we have put out quite a lot of guidance. We will continue to put out a lot of guidance and we will be advising firms who want to set themselves up as ACSPs just because you're an ACSP doesn't mean you have to offer verification now. We expect people will establish themselves as ACSPs in order to file accounts for their clients, but that doesn't necessarily mean they're going to take on new clients just for the purposes of verification. So the website for more detail and start date November 18th. November 18th, anyone, yes, who wants to register and establish a new company from November 18th will be required to immediately provide their identity verification either to companies house or through the ACSP. For those already established on the register, they have a one year grace period, but it's probably best to do it sooner rather than later. That's really helpful. Thanks very much, Mike. Thank you. We're going to wrap up with a look at the inner workings of the OBR. Public finance expert and advisor to ICAW, Martin Wheatcroft is with me. Hello, Martin. Hello, how are you, fellow? Good, welcome back. Shall we start at the beginning? When and why was the OBR set up? The traditional thing when a new government comes in is to blame their predecessors for everything that's gone wrong. And that was what happened when George Osborne came as chancellor with the coalition government in 2010. And the perceived weaknesses in the previous governments, management of the public finances led him to introduce the Office for Budget Responsibility. That's partly to address concerns about the temptation the Treasury might have to massage fiscal forecasts to get the right answer. But also it's best practice internationally. Many other countries have a similar body. So specifically, what is it supposed to do? Well, specifically, what it does is it prepares independently the fiscal, economic and fiscal forecasts for the government. And that's what the Treasury then uses for the budget. And that independent preparation means that the debt markets in particular have some confidence that the forecasts are not being filled about with by the Treasury. I mean, it feels to me, I don't know whether this is accurate. It feels to me a little bit shadowy. I don't even know how many people work there. How big is the organisation? It's a relatively small organisation. There's about 50 people in total of about 30 to 40 are economists and forecasters who actually do the work of the ABR. It's headed by a five-person board comprising three executive board members, what the so-called Budget Responsibility Committee. Who appoints them? So the government appoints them, but it's a fairly robust process. And everybody who's been appointed is a sufficiently independent individual. A couple of them have worked for the Treasury in the past, but they're now independent and seem to be independent. Okay, so that's a perfectly robust process. Is it a fixed-term role? Yes. So Richard Hughes, for example, the chair has a fixed five-year term. He's just been appointed this year for a second five-year term. And that's similar to his predecessor, Sir Robert Chote, who was served for 10 years as the first OBR chair. Okay, now we often sloppily talk about OBR forecasts, but it actually is projections, isn't it? Yes, yes. So technically, they don't do forecasts. They don't try and predict what will happen. What they do is put together an estimate of how the world might look if things happen as we expect them as of the date of their projections. So they start by updating their model for the economy. They look at trends for inflation, interest rates, employment, productivity, migration, international trade, all those good stuff. And then they turn that into a projection for tax receipts and welfare payments based on the current welfare rules and tax rules and the level of interest, debt interest as well. And then over the way that with the previously announced government spending plans. So that's based on the three-year spending review that happened earlier this year. Okay, so the intention is they bring all this specific data in and then they attempt to look at real-world outcomes taking into account perhaps unintended consequences. Yes, yes. So every six months or so they update the projections for what's happened in the real world and altering views of the future. And then they turn the handle again. And then the government then gives them their plans for that they're going to put into the fiscal event. So for example, tax rises, spending changes, all those sort of things. And the OBR makes an estimate of what the economic impact of those because if you increase taxes, you might get mechanically an extra bit of tax. But there might be an economic consequence to that that means that you don't get the full amount. And that's what we saw with the national insurance rise earlier this year. We've got less because employers cut back on staffing and turn into not the full amount of tax receipts that we might have otherwise expected. Okay, so they effectively run their first draft past the government. Then the government tells them their spending plans and then they fact that the spending plans in. Is there a cut-off point? Is there a timeframe beyond which they have to know? Obviously they need to know what they're spending plans on. Well, I mean, it's actually relatively close to the budget. It's a few days before, about a week before, but they're continuing to turn that handle because depending on what the OBR's view is on things, the Treasury will either go back and say, we think it might be better than this or they'll say, ah, we still need a bit more money. So we're going to come up with idea number two or three or 50. So there is a bit of horse trading. So there's a bit of horse trading in that iterative process of preparing the projections. But it's also, you know, quite a rigorous process because the OBR is designed to provide some rigor to the process. But how accurate do these forecasts tend to be? There's two ways of looking that. So my personal view is the projections are always perfect and accurate. The problem is reality is usually wrong. According to the facts as they have them, their projections are excellent. Their projections are excellent. Of course, it's not quite as perfect as that. And the OBR is by no means prescient and able to predict the future very well. But they do their best. But of course, reality comes along. And so, for example, talking about the employee national insurance, the economic damage that's caused has been much higher than the OBR expected. And so that's one of the challenges the OBR will come in and they will reassess when they're reassessing their projections at the moment. Now, obviously, these projections, they really matter, don't they, for a wide variety of sectors, not just Westminster than Chancellor. But what is your sense of how Westminster views the OBR? I think the key thing just to step back a little bit and remember that the OBR is just one part of a bigger system that consists of fiscal rules that the Chancellor has and a whole fiscal responsibility framework that George Osborne introduced at the time. And that overall framework means that the OBR has a couple of different effects. One, it's particularly constraining on Chancellor's means that they are constrained in their choices. And that was part of the design. The problem is that it was very popular with George Osborne when he was Chancellor, but less so with many of his successors. And less so now, as we understand it. Yes. And so, you've had an evolution. So the OBR, I think, is a pretty operationally independent institution that is respected, particularly in the economics world and by many people, but disliked by sort of back benches and cabinet members other than the Chancellor, because it's a tool by which the Chancellor says no. That's quite important for Chancellor's. I mean, it is a good tool for them to say no. The problem is that it also says no to the Chancellor as well. Well, yes, which brings us to, you know, there has been talk, hasn't there, of Rachel Rees trying to find a way to minimize the impact of the OBR on her decision making? Yes. Although, I mean, as I said, it's probably the other parts of the framework that are more challenging for her, her fiscal rules that she set. And it's insisting that she's going to stand by the real constraints. And the fact that she's left herself very little headroom. So the problem in the spring statement was that she had such little headroom, relatively small changes meant that she lost her headroom. And then she had to fiddle around with the numbers a little bit to make the spreadsheet work. But you can see how, from her point of view, it might be quite helpful if the OBR had to, for example, factor in government's growth measures or perhaps didn't do quite so many forecasts every year. Yes. I think one of the challenges is that the OBR does factor in the growth measures. The real problem is that they tend to factor them in after the five-year time horizon. So you do get a boost of growth. But of course, most of these measures that the government is bringing in are not instantaneous. They take while to flow through to the economy. So the OBR scoring of them is not in time to help the fiscal rules, which are based on a five- or four-year time horizon. She's not loon, is she? The New Economics Foundation has gone further. They've called for the forecasting to be taken back in-house, back into the Treasury. I mean, what do you make of that? I think there's a real concern there around the confidence of debt markets, because now they've been introduced to independent forecasting. There would be a real concern about the motivation for doing that and the temptation of the Treasury just to tweak the forecast a little bit, obviously for very good reasons. But that's a slippery slope. And so, whilst I wouldn't say the OBR is sacrosanct, and there are scenarios in which you could see a different arrangement, but it's difficult to see how you get from there to here, particularly at a time when debt markets are quite sensitive to what's going on with the public finances, because they aren't in a great shape. So, it speaks to probity. It speaks to reliability. Indeed, yes. Would you like to see it just left alone? I'm sort of in two minds here, because I think the process is quite good, but the challenge is that in the current context, the government's in a hole, and how it gets out of that hole is quite difficult, and the process and the fiscal rules in particular are making it difficult for the Chancellor to get out of that hole. But you wouldn't argue that she should tweak the OBR in order to assist herself? I think the civil service phrase is, "That would be very brave, Minister." And so, I would not recommend it as a tactical thing, and I think it would be something that you'd need to come up with a replacement system that gave confidence to the markets, because they are nervous. Could have been a lot of unintended consequences there. Martin, thank you very much, as always. Next time on the podcast, we'll be dissecting the Financial Reporting Council's new guidance on using AI in Audit. We'll have an FRC guest to talk us through it, along with two of the senior auditors who fared into that guidance. Over on our sister podcast, The Tax Track, the team are taking another look at making tax digital with chartered accountant and MTD expert Rebecca Beniworth. Just how well prepared is the profession right now? If that's your area, you can find the podcast on any app. Thanks for listening. [Music]

Podcast Summary

Key Points:

  1. Discussion on the value and limitations of large language models (LLMs) for spreadsheet tasks.
  2. Companies House ID verification requirements for directors and persons of significant control.
  3. Distinction between AML supervision and ACSP verification causing confusion.

Summary:

The transcription covers Accountancy Insights discussing the value and limitations of large language models (LLMs) for spreadsheet tasks. The conversation delves into the research conducted by Simon Thorne on LLMs and their performance in various spreadsheet tasks. It highlights the inconsistency in LLMs' abilities and their limitations in more complex scenarios.

The discussion also touches on the Companies House ID verification requirements for directors and persons of significant control, emphasizing the need for compliance and the distinction between AML supervision and ACSP verification. The conversation reveals potential challenges with overseas directors and the importance of thorough ID verification. Overall, the transcription provides insights into the evolving landscape of technology in accounting and the regulatory requirements facing professionals in the field.

FAQs

Accountants should use large language models for time-saving tasks, but must validate and verify the output to ensure accuracy and completeness.

Personally identifiable information should be avoided in large language models to prevent potential training on sensitive data and privacy breaches.

Accountants should remind clients to complete ID verification, especially for directors, PSCs, and LLP members, to comply with Companies House requirements.

Accounting firms may need to supervise clients' ID verification and consider becoming Authorized Corporate Service Providers in the future.

Accounting firms may face confusion between AML supervision and ACSP verification requirements, especially when dealing with overseas directors and different document standards.

Chat with AI

Loading...

Pro features

Go deeper with this episode

Unlock creator-grade tools that turn any transcript into show notes and subtitle files.