Round stones lays on a grass. Blue sky on a background

Introducing the English Grammar Profile #1 – Building the profile

Share on Facebook Share on LinkedIn Share on Twitter

In the first of two posts, Geraldine Mark and Anne O’Keeffe introduce the English Grammar Profile and explain how it was created.

When we were approached by Cambridge five years ago to research and write the English Grammar Profile, it seemed like an exciting and daunting prospect – to map out each and every aspect of learner grammar at each level. It was a task which would take us into our old age, we thought. Our job was to find out what grammar learners can competently use at each level of the CEFR. The important thing was not to produce a list of what grammar we thought they could use but to look at what the evidence shows about what they are actually using.

So where did we get our evidence from?

We used the Cambridge Learner Corpus (CLC), which is a sub-corpus of the Cambridge English Corpus (CEC). The CLC is a 55-million-word corpus consisting of over 200,000 writing scripts of all levels, taken from Cambridge exams. The learners represented in this data have between them over 140 first languages, from more than 200 countries. It’s quite a body of data! Not only that, but of the 55 million words, 32 million have been tagged for errors. In other words, every time there is a grammatical error, it has been marked up with some code, so that it is easy then to find all of the recurring errors, and to see at what level they are being made, and so on.

How did we create the profiles?

We started with a list of the main grammar points that you can find in coursebooks and grammar books – a list that teachers would typically expect to teach (consisting of tenses, words classes, types of clauses, modals, etc.). We then broke these down into subcategories depending on the items: so, for example, for pronouns the subcategories would be subject, object, possessive, etc.). Our challenge was to produce an item-by-item description of all of the things that learners could do – a series of ‘can-do statements’ describing learners’ use of grammar. We were very fortunate to be able to use SketchEngine – an online corpus analysis tool which allowed us to ask questions of the data. This allowed us to ask questions about CEFR level, pass/fail results, nationality, first language, contexts, task types, exam, year of exam, candidate age…

funnel copy

Our methodology

For our methodology, we established a type of ‘filtering process’. What we did was put a structure into the ‘filter’ and then analyse the results at each stage so as to find out when learners could competently use a structure. For each item we investigated, we started with a structure at A1 and looked to see if there were any results. If there weren’t, we moved to the next level of the CEFR. The answer to each question had to be ‘yes’ before the structure was ‘allowed’ into the next stage of the process. If the answer was ‘no’ at any point, we moved to the next CEFR level and started the filtering process again.
If the answer to every question was ‘yes’ we wrote a ‘can-do’ description of the evidence.
Eventually, by using this ‘filtration method’ we arrived at a level where:

  • an item was frequent enough in terms of how often it was used.
  • an item had enough correct uses.
  • these correct uses were spread across enough different learners
  • these correct uses were spread across enough different first language families
  • these correct uses were spread across enough different contexts of use
  • these correct uses were not affected by the task (e.g. the learners were not taking the grammatical pattern from the rubric)

We went through these stages of analysis for all of the grammar items. That’s why it took us four years! The result is the more than 1200 descriptions of the English Grammar Profile which Cambridge put into an online format that is freely available and searchable as a database.


To explore the English Grammar Profile for yourself, click here.