I once received a sample from a potential supplier in Colombia. The roast looked beautiful. Uniform, even, a nice medium-brown. The dry fragrance was intoxicating—blueberry, dark chocolate, a hint of jasmine. I was ready to order a container. Then I cupped it. The acidity was sharp and sour, not bright. The body was thin. The finish had a papery, stale note that lingered unpleasantly. The roast was masking a mediocre green bean. The supplier had a talented roaster, but they didn't have great coffee. If I had bought based on the roast alone, I'd have overpaid for a lot that couldn't deliver in the cup.
Evaluating sample roasts from a potential supplier requires separating the roast quality from the green bean quality. A skilled roaster can make a mediocre bean look promising. A poor roast can make a great bean taste flat. The buyer's job is to see through the roast to the bean underneath—assessing the green bean's intrinsic quality, its potential, and the supplier's consistency. This means evaluating the green sample first, controlling for roast variables, and cupping systematically, not impressionistically.
I've been on both sides of this equation. I send samples from Baoshan to buyers around the world, and I evaluate samples from other origins. Here's the system I've developed over years of being fooled, surprised, and occasionally vindicated.
What Should You Look for in the Green Sample Before Roasting?
Before a single bean touches the roaster, the green sample tells you more than the roast ever will. A roast can hide defects. A green bean cannot. The physical evidence is right there on the grading mat—the size, the shape, the color, the smell, the presence of defects. Skipping the green evaluation is like buying a used car without opening the hood. The paint might be shiny, but you have no idea what's inside.
Green sample evaluation should focus on four objective criteria: moisture content and water activity for stability, screen size and uniformity for roasting consistency, defect count according to SCA grading standards, and sensory evaluation of the raw bean's smell and appearance. A professional supplier provides this data with the sample. A buyer who evaluates it independently builds a complete picture of the green coffee's intrinsic quality before the roasting variables are introduced.

Why Does Moisture Content and Water Activity Matter Before Roasting?
Moisture content and water activity are the vital signs of a green coffee. They tell you whether the bean is biologically stable and how it will behave in the roaster. A supplier who doesn't provide these numbers—or doesn't know them—is not managing their coffee professionally.
I measure moisture content with a calibrated meter on every sample I receive. The target range for specialty green coffee is 10% to 12%. Below 9.5%, the bean is overly dry and may have lost volatile aromatics. It will roast fast and taste flat. Above 12.5%, the bean is at risk for mold and will roast unevenly, with grassy or fermented notes persisting in the cup.
Water activity is the more sophisticated metric. I want to see a reading between 0.53 and 0.60 aW. This is the zone where the bean is stable, safe from microbial growth, and still holding its flavor potential. A water activity above 0.65 is a red flag. Even if the moisture content looks fine, high water activity means the water in the bean is available for mold and bacteria. I've rejected samples on water activity alone, even when the green looked perfect and the moisture number was within range. Water activity predicts storage stability and shelf life. A professional supplier measures it and shares it.
How Do You Conduct a Green Grading for Defects and Uniformity?
Green grading is a systematic process, not a casual glance. I use the SCA Green Arabica Coffee Classification System, even for Robusta samples, because it provides an objective framework. The process is simple but must be followed exactly.
I weigh out a 350-gram sample. I spread it on a black grading mat under good light. I sort through the entire sample, picking out every defect. Primary defects—full black beans, full sour beans, large stones, large sticks—count as one full defect each. Secondary defects—partial blacks, partial sours, broken beans, insect damage, small sticks—count as fractions. The SCA specialty threshold is zero primary defects and fewer than five secondary defects per 350 grams.
Uniformity matters as much as the defect count. I measure the screen size distribution. A sample with beans ranging from screen 14 to screen 18 will roast unevenly. The smaller beans will scorch before the larger beans are developed. A professional lot is screened to a narrow range, typically screen 16 and above for specialty Arabica. Uniform bean size means predictable roasting behavior. Unpredictable roasting means unpredictable cupping. I want predictability.
I also smell the green sample. Clean green coffee smells faintly sweet and vegetal, like fresh hay or green tea. A musty, smoky, or chemical smell indicates a processing or storage problem. The nose knows. Trust it.
How Can You Control for Roast Variables When Comparing Suppliers?
Comparing samples from different suppliers is only fair if the roast variables are controlled. If Supplier A's sample is roasted light and Supplier B's is roasted dark, you are comparing roast profiles, not coffee quality. The darker roast will taste more developed, more chocolatey, less acidic—regardless of the bean underneath. To compare suppliers fairly, you must remove the roast as a variable as much as possible.
Controlling for roast variables means using a standardized sample roasting protocol for all incoming samples. The protocol specifies the batch size, charge temperature, development time, and end temperature or color target. Roasting all samples to the same Agtron number or color range ensures that the cupping comparison reveals differences in the green coffee, not differences in the roast degree. A log of each roast is essential for identifying outliers and troubleshooting off-flavors that may be roast-derived rather than bean-derived.

What Is a Standardized Sample Roasting Protocol?
A standardized protocol is not a single recipe. Different beans, especially from different origins and processing methods, roast differently. A washed Ethiopian Arabica and a natural Brazilian Arabica will need slightly different approaches to reach the same development. The protocol is a target, not a straightjacket.
My standard sample roast targets an Agtron ground color of 55 to 60, which is a light-medium roast suitable for cupping. I use a 100-gram batch size in a calibrated sample roaster. I preheat the roaster to a consistent charge temperature. I aim for a first crack time between 7 and 9 minutes, depending on the bean density. I target a development time—the time from first crack to drop—of 1:30 to 2:00 minutes. This is enough development to avoid grassy, under-roasted notes without introducing roast character that masks origin flavor.
I log every roast. The charge temperature, the turn-around point, the rate of rise at each minute, the first crack time, the drop temperature, and the total roast time. If a sample cups poorly, I check the roast log first. Was the development time too short? Did the rate of rise stall or flick? A defect in the roast curve is a controllable error. I never judge a green sample based on a roast I know was flawed. I roast it again.
Why Should You Use a Reference Sample in Every Cupping Session?
A reference sample is a known coffee, roasted and cupped many times, that serves as a benchmark for the session. Without a reference, your palate is uncalibrated. You're evaluating samples in a vacuum. With a reference, you're evaluating them against a standard.
I use a washed Central American Arabica that consistently cups at 82 to 83 points as my reference. I roast it fresh for every cupping session, using the same protocol as the evaluation samples. I cup it first, before any of the unknown samples. This calibrates my palate to the session's conditions—the water temperature, the grind size, the ambient room temperature, even my own sensory state that day.
If the reference cups differently than usual—if it tastes flat or bitter—I know something is off with the session, not the samples. I adjust and re-cup. If the reference cups normally, I have a baseline for scoring the unknown samples. A sample that scores 85 against a reference that I know is 82 is a genuinely good coffee. A sample that scores 85 but the reference also scored 85 today is just a normal coffee evaluated on a generous day. The reference keeps you honest.
What Are the Key Sensory Markers of Quality in a Sample Roast?
The sensory evaluation is the heart of the sample assessment. It's where all the data—the green grading, the moisture readings, the roast log—converges in a single, subjective, irreplaceable experience. The coffee either tastes good or it doesn't. But "good" is not a useful evaluation. You need to break "good" down into specific, observable sensory markers that you can compare across samples and across sessions.
The key sensory markers of quality in a sample roast are cleanliness of cup, clarity of flavor, balanced acidity, body appropriate to the origin and processing method, and a clean, persistent finish without astringency or off-notes. These markers should be evaluated systematically using a cupping form, not impressionistically. The best samples demonstrate not just intensity of positive attributes, but the absence of negative ones.

How Do You Distinguish Origin Character from Roast Character?
This is the hardest skill in cupping. Roast character is the flavor contribution of the roasting process itself. Origin character is the flavor contribution of the bean—the variety, the terroir, the processing. They overlap, and untangling them takes practice.
Roast character tends to express as caramelization, toast, dark chocolate, and sometimes a pleasant bitterness. These are the Maillard reaction products and the caramelization of sugars. In a well-developed light roast, these notes are subtle and supportive, not dominant. In a darker roast or a roast with too much development, the roast character overwhelms the origin character. You taste the roast, not the coffee.
Origin character expresses as fruit, floral, herbal, and spice notes. A washed Yirgacheffe might show lemon and jasmine. A natural Brazilian might show blueberry and cocoa. A washed Yunnan Arabica from our farm might show red apple, brown sugar, and a hint of black tea. These notes are delicate. They are easily buried by roast character.
If a sample tastes generically "coffee-like"—roasty, a little sweet, a little bitter, but without any specific, identifiable flavor—the roast is probably too dark or too developed. Ask for a lighter roast. If the supplier cannot or will not provide a lighter roast, be suspicious. They may be hiding a bland or defective green bean behind a wall of roast flavor. A confident supplier roasts light enough to let the origin speak.
What Off-Flavors Indicate Green Bean Defects Rather Than Roast Problems?
Some off-flavors are roast problems. Grassy, hay-like notes usually mean under-development—the roast was too light or too fast. Scorched, smoky, ashy notes usually mean the roast was too hot or the drum was dirty. These are correctable. You can ask for a re-roast.
Other off-flavors are green bean defects, and they cannot be roasted away. A fermenty, vinegary sourness indicates over-fermentation during processing. A musty, moldy, basement-like flavor indicates moisture damage or mold growth. A phenolic, medicinal, band-aid-like flavor indicates a serious microbial contamination during processing or storage. A baggy, papery, woody flavor indicates age—the coffee is past crop and has absorbed storage container flavors.
These defects are not subtle. Once you've tasted them, you recognize them instantly. A supplier who sends a sample with these defects either doesn't cup their own coffee before sending it, or doesn't care. Neither is acceptable. I discard these samples and move on. There are too many good suppliers in the world to spend time on ones who send defective samples.
How Should You Compare Multiple Supplier Samples Objectively?
Evaluating one sample is straightforward. Evaluating five samples from five different suppliers and choosing the best one is hard. The palate fatigues. Bias creeps in. You want the cheapest sample to win, or the one from the supplier you already have a relationship with, or the one with the prettiest packaging. The only way to make a fair comparison is to remove the identifying information and use a structured scoring system.
Comparing multiple supplier samples objectively requires a blind cupping setup with random coding, a standardized scoring sheet, and a structured session that limits palate fatigue. Samples should be evaluated across multiple sessions, not just one, to account for day-to-day sensory variation. The final decision should integrate the cupping scores with the green evaluation data, the price, and the supplier's reliability indicators, not rely on the cupping alone.

How Do You Set Up a Blind Cupping for Fair Comparison?
A blind cupping is simple to set up and dramatically improves the fairness of your evaluation. Have someone who is not participating in the cupping assign random three-digit codes to each sample. The coder keeps the key. The cuppers do not know which sample is which.
The cupping bowls are labeled only with the codes. The samples are roasted to the same color target, ground to the same grind size, and brewed with the same water at the same temperature. The only variable is the coffee itself. This is not a perfect scientific experiment, but it is a fair one.
I cup each sample individually, scoring each attribute on the SCA cupping form—fragrance/aroma, flavor, aftertaste, acidity, body, balance, uniformity, clean cup, sweetness, and overall. I write descriptive notes, not just numbers. "Black cherry, milk chocolate, clean finish" is more useful than "8.5." After scoring all samples, I rank them by total score and by overall impression. Then the codes are revealed. The reveal is often surprising. A sample I was prepared to love because of a beautiful offer sheet sometimes finishes third. A sample from an unknown supplier sometimes wins. Blind cupping forces you to trust your palate, not your preconceptions.
Why Should You Re-Cup Finalist Samples Before Making a Decision?
A single cupping session is a snapshot, not a definitive assessment. Your palate is different on Tuesday than it was on Monday. The roast might have been off by a degree. The water might have been slightly cooler. The sample might have needed another day of rest after roasting. For a sourcing decision that commits thousands of dollars, a single cupping is not enough.
I re-cup my top two or three samples at least once, and preferably twice, on different days. I may roast fresh batches for each session. If the scores and the impressions are consistent across sessions, I have confidence in the evaluation. If a sample scores 86 on Monday and 83 on Wednesday, something is unstable. It could be the coffee, or it could be my palate, but either way, I need to understand the discrepancy before I commit.
I also encourage suppliers to send pre-shipment samples from the actual lot that will fill my container, not just a representative sample. The pre-shipment sample is the final exam. It's the coffee that will arrive at my warehouse. I cup it against the original offer sample. If the two cup differently, I ask why. A professional supplier can explain any variation—different roast dates, different portion of the lot, different resting time. A supplier who can't explain the difference is a risk I don't take.
Conclusion
Evaluating sample roasts is a skill that separates professional buyers from hopeful ones. It requires looking past the roast to the green bean underneath, controlling for variables so the comparison is fair, and using structured sensory evaluation rather than gut feeling. The best buyers combine objective data—moisture readings, defect counts, Agtron numbers—with subjective but disciplined cupping, and they verify their impressions across multiple sessions.
The suppliers worth building a relationship with are the ones who make this evaluation easy. They send green samples with full data sheets. They roast light enough to let the origin character show. They welcome pre-shipment sample approvals and don't make excuses when something is off. At BeanofCoffee, I treat every sample we send as a test of our professionalism. The green grading data, the water activity reading, the cupping notes, the roast recommendation—it's all in the package. If you're evaluating suppliers and want to put our Baoshan coffees through your protocol, reach out to Cathy Cai at cathy@beanofcoffee.com. She'll send you a sample set with everything you need to make a fair, informed decision. Cup it blind. Cup it twice. We're confident in what you'll find.