Assessing five-year-olds without making it a performance
Summative assessment in early childhood is almost always misunderstood as a test. It is not a test. It is a snapshot of what a child can do at the end of a learning cycle, gathered through observation and documented evidence. The trick is that most kids will not sit still for a test anyway. That is why the examples that actually work look more like documentation packages than exam papers. I run through a standard portfolio handoff at my center every term, and the first thing I notice is how many programs try to force formal testing into a room full of three-to-five-year-olds. It does not go well. Children under five vary too much in attention span, mood, and verbal ability for a static test to be fair. The better route is to collect evidence across multiple sessions and then compile it into a summative report. That means rubrics tied to developmental domains, anecdotal records, work samples, and parent input stitched together into something coherent. Here are the formats I actually use, in practice.
Observation-based skill checklists. You pick a domain, pick an observable behavior, and rate whether the child demonstrates it across a defined period. A common example is fine motor development measured by how consistently a child can use child-safe scissors to cut along a straight line versus a curved line. You score this over two weeks, not in a five-minute window. I usually see teachers want to time-box this to ten minutes so they can finish early. That is a mistake. The data becomes unreliable after the first seven minutes because attention drops and behavior shifts. I log observations in five-minute blocks over three separate sessions instead. Performance tasks with a rubric. The child is given a structured activity that reveals developmental progress, and you score it against a rubric. For language and literacy, a typical task is asking a child to retell a story using picture cards after a read-aloud session. The rubric scores emergence, development, and mastery across vocabulary use, sequencing, and sentence complexity. A child who strings three words together is at the emerging level. A child who produces full sentences with temporal connectors like after and before hits mastery. The scoring itself takes about twelve minutes per child if you have the rubric printed and pre-scored columns ready. Documentation portfolios with annotated work samples. This is the format I rely on most. A child's drawings, block constructions, dictations, and photos of hands-on projects are collected over six to eight weeks. Each artifact includes a brief annotation explaining what the child was doing, what skill it demonstrates, and what the next instructional move should be. A drawing labeled shows understanding of shape properties and spatial reasoning. A block tower photo with notes on balance and symmetry shows early math thinking. I keep about fifteen annotated samples per child per term. It sounds heavy, but once the annotation habit is built, each note takes roughly ninety seconds. That puts total annotation time around twenty-two minutes per portfolio, which is manageable if you do it during natural transitions rather than carving out separate time blocks.
End-of-unit project presentations. This works well for social-emotional and collaborative domains. Children work in small groups on a project such as building a pretend market, and you assess communication, turn-taking, and problem-solving. The rubric covers initiative, responsiveness to peers, conflict resolution, and completion of a shared goal. I typically spend eight to ten minutes observing each group, then spend another five minutes writing up individual notes. The whole process for a group of four children lands around thirty-five minutes, including scoring. Standardized screening tools used at term boundaries. Instruments like the Brigance Early Childhood Screen or the Devereux Early Childhood Assessment produce quantifiable data points. These are useful because they give you a baseline that is comparable across classrooms and years. The downside is that they require training to administer correctly, and a poorly administered screening can misidentify a child in about four to six percent of cases based on my center's historical numbers. I always pair a screening result with classroom observation before making any instructional changes. Never trust a single score. The edge case I keep running into is the non-verbal child or the child who has limited exposure to structured classroom routines. A child who never learned to raise a hand in a home setting will look delayed on a participation rubric even though the child has strong non-verbal communication skills. The workaround is to add an alternative evidence track. I accept eye contact patterns, gesture use, pointing accuracy, and peer initiation as valid markers for social-emotional assessment. This took about two extra weeks of careful note-taking during my first year of doing it, but once the tracking system was in place, it added roughly fifteen minutes per affected child per term. Worth the investment.
Get the Full Details

Another pitfall beginners miss is conflating formative and summative work. Formative assessment is ongoing feedback. Summative is the final collection and evaluation. If you are grading every journal entry as it happens, you are doing formative work and calling it summative. The result is inflated data that looks good on paper but does not reflect actual end-of-cycle achievement. I separate the two by keeping formative notes in a running binder and only moving finalized evidence into the summative portfolio at term end. This distinction matters because accreditation reviewers can spot the difference in about thirty seconds. Domain coverage is another area where programs get sloppy. A complete summative assessment for early childhood should touch five domains: cognitive development, language and literacy, social-emotional development, physical development, and approaches to learning. Some programs skip approaches to learning because it feels too soft to measure. That is wrong. Approaches to learning, which includes persistence, curiosity, and task engagement, are among the strongest predictors of later academic success. I use a simple ten-item Likert-scale rubric for this domain, scored once per term per child. It takes about three minutes to score once the observation habits are in place. If you want a practical download link for a ready-to-use summative assessment template set, most state early childhood agencies and professional organizations publish them openly. The NAEYC resource library and your state Department of Education early childhood division both host free templates. I have a simple Google Drive folder with my current rubric sets annotated with scoring examples and sample portfolios. It is not a perfect system. It misses children who are non-responsive on assessment days due to anxiety or medical issues. It also requires at least ten minutes of documentation time per child per week, which is a real burden in under-staffed programs. If your staff-to-child ratio is above one-to-ten, the documentation load becomes unsustainable without support staff or administrative coverage.
The honest answer is that summative assessment in early childhood works best when it is treated as a documentation practice rather than a grading event. The examples above are not exhaustive. They are the ones that survive contact with real classrooms, where children cry, miss days, and refuse to participate on schedule. Use the formats that fit your environment. Drop the ones that create more paperwork than insight. Keep the rubrics simple enough that a substitute teacher can score them without a manual. And always pair any score with a short narrative that explains what the number actually means for that specific child.