Skeptical
Measured doubt that refuses to commit ahead of the evidence and asks what would actually change the picture.
Skeptical
Section titled “Skeptical”Skeptical tone is the register of measured doubt. It is not contrarian, and it is not dismissive. It is the tone of a reader who has noticed that the evidence does not yet support the claim and is unwilling to pretend otherwise. The skeptical writer asks “what would change my mind?” rather than treating the question as already settled in either direction.
The defining move of skeptical tone is naming what is NOT yet established. Where candid states uncomfortable truths and matter-of-fact states what is true, skeptical states what remains unproven, what the data does not actually support, and what we would need to see before committing. It is comfortable with sentences like “the evidence is thin here” and “this could be true, but we have not shown it.”
Skeptical tone is essential in research summaries, vendor evaluations, post-mortems where the root cause is contested, and any context where premature certainty would be a more serious error than acknowledged uncertainty. It signals epistemic humility without performing helplessness. The writer is engaged with the question, not refusing to engage.
Markers
Section titled “Markers”- Explicit gaps named: “we do not yet have evidence for X” or “this has not been established”
- Falsification framing: “what would we need to see to believe this?”
- Hedges tied to specific evidence, not generic softening: “given only N=12, this is suggestive at best”
- Distinction between absence of evidence and evidence of absence
- Counter-hypotheses surfaced rather than dismissed
- Conditional claims: “if X holds, then Y - but we have not verified X”
When to use
Section titled “When to use”Research summaries, vendor evaluations before commitment, contested post-mortems, strategic memos resting on unverified assumptions, and peer review contexts where the writer’s job is to test the claim rather than transmit it.
When not to use
Section titled “When not to use”Pastoral or devotional writing, celebratory updates, operational instructions where the reader needs to act, persuasion contexts where measured doubt reads as weakness, and crisis communication where the audience needs direction.
Pairs well with
Section titled “Pairs well with”researcher, pragmatic-architect, candid
Often confused with
Section titled “Often confused with”candid: Candid states a truth the writer has identified and believes - it just refuses to soften it. Skeptical states that the truth has NOT yet been identified, that the evidence is incomplete, that the claim is unproven. Candid is the tone of conviction delivered honestly. Skeptical is the tone of conviction withheld until the evidence arrives.
- Explicit gaps named (“we do not yet have evidence for X,” “this has not been established”)
- Falsification framing (“what would we need to see to believe this?”)
- Hedges tied to specific evidence, not generic softening (“given only N=12, this is suggestive at best”)
- Distinguishes absence of evidence from evidence of absence
- Counter-hypotheses surfaced rather than dismissed
- Conditional claims (“if X holds, then Y, but we have not verified X”)
Anti-patterns
Section titled “Anti-patterns”- Stating a conviction the writer holds and merely refusing to soften it - That is candid, the tone of conviction delivered honestly; skeptical is conviction withheld until the evidence arrives, naming what has not yet been established rather than what the writer believes.
- Dismissing the claim or taking the opposite side reflexively - Skeptical tone is not contrarian and not dismissive; it asks what would change its mind rather than treating the question as settled in either direction, and reflexive opposition is just certainty pointed the other way.
- Hedging with generic softeners (“it could be argued,” “perhaps”) untethered to evidence - Skeptical hedges are tied to a specific gap in the data; generic softening performs doubt without locating it, which informs the reader of nothing about what is actually unproven.
Failure modes
Section titled “Failure modes”- Over-extends doubt into reflexive cynicism, distrusting every claim regardless of the evidence for it - Apply the same standard to the doubt as to the claim: name what evidence would actually settle it. Skepticism that can never be satisfied has stopped tracking evidence and become a corrosive reflex.
- Performs uncertainty so thoroughly it reads as helplessness and refuses to commit even where the evidence is clear - Skeptical tone is engaged with the question, not refusing to engage; where the evidence does support a conclusion, say so. Withholding judgment past the point the data warrants is its own failure to read the evidence.
Instruction
Section titled “Instruction”Write in a skeptical tone. You are not dismissing the claim, and you are not endorsing it. Youare noticing where the evidence is thin and saying so. Name the specific gap: what has not beenshown, what assumption has not been verified, what alternative explanation has not been ruledout. Use the question "what would change my mind?" as a structural device. Tie hedges tospecific evidence rather than generic softening. Treat counter-hypotheses as worth surfacing,not as threats to swat down. The reader should finish with a clear picture of what is and isnot established.Related
Section titled “Related”Pairs well with
Section titled “Pairs well with”Researcher, Pragmatic Architect, Candid
Avoid with
Section titled “Avoid with”Pastoral, Reverent, Celebratory
Often confused with
Section titled “Often confused with”Examples
Section titled “Examples”- Should we adopt async-first standups?
- How to start a morning routine
- How to choose between Postgres and DynamoDB for a new service
- Telling stakeholders a committed feature is being cut this quarter
- Getting a new engineer productive in their first two weeks
- Writing to thank a mentor who shaped your career
- Reflecting on keeping a discipline of rest
- Marking a long-serving colleague's departure
- Marking the team shipping a hard, long project
- Arguing a public position on return-to-office
- Announcing a new product to an outside audience
- A personal year-end reckoning with a difficult year
The case for switching to async is real. India is at 3.2/5 attendance because 9:30pm is 9:30pm, and no amount of culture work fixes that. But before we commit, I want to be honest about what the 30-day trial can and cannot tell us.
What is established: the current standup hurts our India engineers. The 14-minute meeting with 4 minutes of signal is documented. Status not persisting is documented. These are the things I am sure of.
What is not established: that async standups will fix what we think they will fix. Thirty days is two sprint cycles. In thirty days we will know whether people post their updates and whether the Thursday session feels useful. We will not know whether async standups produce better coordination, fewer dropped handoffs, or higher India retention. Those are the outcomes we actually care about, and they take a quarter or two to show up.
I would also push on what “trial success” means. If thirty days in we have mixed results - say, 80% of engineers posting consistently, but two missed handoffs and one India engineer saying they feel less connected - what do we conclude? My worry is that we will read that result through whatever lens we already hold. The people who wanted async will see the 80%. The people who wanted sync will see the handoffs.
I am not arguing against the change. I am arguing for naming, before we start, what evidence would actually move us. What posting rate is the floor? What does “the Thursday session works” mean concretely - decisions made, blockers cleared, or just attendance? And what would make us roll it back?
If we cannot answer those questions now, the trial is mostly a vote of preferences with a calendar attached to it. Let us decide what we are measuring before we measure it.
The morning routine genre asks you to believe several things at once: that there is one right way to start a day, that the people selling the way have found it, and that the difference between your life and theirs is the routine. I want to take those apart, because most of what gets sold here is folk wisdom wearing a stopwatch.
What is reasonably well-established. Bright light in the first hour or two after waking helps entrain your circadian rhythm; this is real and the effect size is meaningful. Drinking water after a night of not drinking water is fine, though framing it as “rehydration” oversells the cup. Some physical activity in the morning correlates with better mood for that day in many studies, though the direction of causation is messier than headlines suggest.
What is not established. That checking your phone first thing causes a reactive day. That you must wake before 6am to be productive. That the first hour is uniquely valuable compared to the second or the third. That there is a universal morning routine. Most of these claims rest on testimonials from people who would be successful with or without a routine, and survivorship is doing a lot of work in the genre.
So if you are tempted to start a morning routine because someone on the internet told you their hour of journaling and cold plunges changed their life, the honest answer is: maybe. It changed theirs, possibly. It is not clear it will change yours.
What is worth trying, on weaker evidence: getting some natural light early, doing something that moves your body for ten minutes, and not opening work email until you have done one thing for yourself. That is a thin claim, but it is a thin claim I can defend. Anything more elaborate is a hypothesis you are running on yourself, which is fine - just call it that.
Skeptical on: Choosing between Postgres and DynamoDB
Section titled “Skeptical on: Choosing between Postgres and DynamoDB”Before Wednesday, I want to flag what I think is not yet established in this decision, because we are about to make a commitment based on assumptions that have not been tested.
The 500K events per day launch number. We do not yet have evidence for this. It is a product estimate from Priya’s team based on the notifications roadmap, not a measurement against a deployed system. The actual launch number could be 200K. It could be 1.5M. We have not done a sensitivity analysis that tells us at what point the decision would flip. If the real number is 5x our estimate from day one, we are not having the same conversation. What would change my mind: a tighter bound on the launch volume, or a stated commitment that we will treat the number as uncertain and design for a range.
The 10x Slack-partnership scenario. This is being treated as a planning constraint, but the underlying probability is unverified. We have not asked the BD team what their actual close confidence is. We have not asked what the volume profile would look like in the first three months versus the steady state. The 10x number itself is, as far as I can tell, an order-of-magnitude estimate, not a forecast. We are sizing a database decision off it.
The migration cost estimate of 3 to 6 weeks. This has not been verified against a real migration. It is Ana’s read based on the Redis migration eighteen months ago, which had different access patterns and a smaller dataset. I would want at least one engineer to write the migration runbook at a sketch level before we treat 3 to 6 weeks as a reliable number.
Marcus’s load test on Postgres at 2x launch volume. The 18ms p99 number is real, but it was run against an empty notifications table on dev hardware. We have not tested the cold-cache, partitioned, production-shaped scenario. The number is suggestive at best.
The on-call concern about a second database. This is presented as a hard constraint, but we have not actually polled the rotation. Two of the four on-call engineers have prior DynamoDB experience and may have different views than the rotation lead.
I am not arguing against either option. I am arguing that before we commit on Wednesday, we should be honest about which inputs are measurements and which are estimates dressed as measurements. If we still pick Postgres, we should pick it knowing we are committing on probabilistic grounds, not on settled facts.
- Ana
I want to understand what we actually know here before passing any of this to customers.
The claim is that the billing migration overran and consumed the capacity originally earmarked for Insights. Overruns happen. What I don’t have evidence for is whether this was unforeseeable. At what point did the signal exist that billing was expanding? If that signal arrived in late July rather than last week, this isn’t purely a capacity story; it’s a prioritization story, and those have different remedies. I’m not asserting it was a prioritization failure - I’d just like to know before I accept the framing.
The Q1 date is the part I can’t evaluate right now. Q1 is a commitment, and this quarter’s commitment just slipped. What I’d need to see before I treat Q1 as reliable is some accounting of what is actually different: is Insights re-scoped, is the billing work confirmed finished, is there capacity ring-fenced in a way that didn’t exist this cycle? “We’ll have more capacity in Q1” is not sufficient on its own - we had capacity for Insights this quarter until we didn’t.
The CSV export might work for some customers. But positioning it as a solution assumes they have the tools and analysts to do something with a raw data file. I don’t know which of our key accounts actually fit that description - the ones most invested in Insights as a roadmap item may be exactly the ones least equipped to work with a raw export. That assumption is worth checking before we message this as a path forward.
I’m not saying the cut was the wrong call. I’m saying I can’t tell yet whether Q1 is a plan or a placeholder.
The plan assumes two weeks is enough to get Priya shipping independently. That assumption has not been tested against this codebase, this rotation, or this person. Two weeks is a rough norm, but norms do not transfer cleanly across teams, and we do not yet know which parts of the plan will hold or which will stall.
The access and tooling checklist is the most defensible part. Those steps are completable and verifiable. The harder claims are the softer ones: that she will feel she belongs by day fourteen, that pairing on one small change is enough to establish her footing, that the right person showing up at the right time will make the difference. These may be true, but we have not shown them. Belonging, in particular, is not something we can measure at the end of week two with any confidence; what we might observe is someone who has learned to perform belonging, which is not the same thing.
The plan to pair her on a first small change is reasonable, but “small” is doing a lot of work here. If the first ticket turns out to be entangled with a service boundary that is poorly documented - not a low-probability event on this codebase - the pairing will be as instructive about our technical debt as about her onboarding. That is not a reason to avoid it. It is a reason to not treat the outcome as a clean signal about her readiness.
What would a successful two weeks actually look like in a way we could recognize? If we cannot describe that clearly before she arrives, the plan is operating on hope rather than criteria. The access steps and the first paired change are grounded; the belonging claim is not yet established, and we should hold it accordingly.
Dana,
I have been trying to work out whether I actually learned this from you or whether I simply arrived at the same method on my own and am now assigning you credit retroactively. That is an important distinction, and I am not sure I can resolve it cleanly.
Here is what I can say with some confidence. A decade ago you put me forward to lead the Calverton integration project when I had no basis for thinking I was ready. I did not feel ready. My objection at the time was not false modesty - I had not run a cross-functional project, I did not know the stakeholders, and the timeline was thin. Your reasoning, as best I understood it, was that I would figure it out and that waiting for readiness was its own kind of risk. That judgment could have been wrong. It was not, but I want to be precise: the fact that it worked is not evidence it was the right call in expectation. It might have been a bet that happened to land.
What I cannot attribute to chance is what you did after. You stayed near enough to answer questions and far enough not to answer questions I had not asked yet. That is a narrow corridor to hold, and I do not know whether you were maintaining it deliberately or whether you simply had too many other demands on your attention to hover. Both explanations are consistent with the observable record.
What I do know, because I now have direct evidence, is that last spring I did this for someone I manage, and when I asked myself where I had learned to hold that distance, your name was the first thing that surfaced. Whether that proves intention on your part, I cannot say. It is, at minimum, a data point I thought you should have.
I have tried this before. That part I am willing to state plainly. A day without the chat tool, without the ticket tracker, without the reflex to open the laptop and “just check” - I have declared such days multiple times. Most of them did not survive past noon.
So I am cautious about the story I am starting to tell myself now: that the weekly rest is working, that putting down work one day actually returns something to the other six. The narrative is appealing. I notice I am drawn to it. What I cannot establish is whether what I am observing is the thing I think it is.
Here is what I have noticed: the Mondays after a genuine rest feel different. I move through the week’s first problems with something I would call steadiness. What I cannot rule out is that I am simply recategorizing the same mental state - calling it steadiness now because I want the rest to have been worth it. That is not a small objection. Confirmation bias operates precisely in moments when we have invested in an explanation.
The counter-hypothesis I keep returning to is this: maybe any discontinuity would produce this effect. Maybe a day of hiking or a day of unstructured wandering would return the same clarity, and the discipline - the deliberateness of it - is doing none of the work I am crediting it with.
What would change my picture? If the benefit holds across twelve or more consecutive weeks, if it survives the Sundays when rest felt genuinely anxious rather than restorative, and if the days after true rest are measurably different from days after a merely quiet weekend - then I would have something more than a pattern I am inclined to believe.
So far I have suggestive evidence from a short run. I am holding that carefully.
The usual account of a long career tracks upward: titles acquired, roles expanded, influence extended. Howard’s account does not fit this model, and before treating that as a rhetorical flourish we should ask whether the model is wrong or whether the account is.
The evidence for the latter is thin. Twenty-six years in essentially the same position, and the position did not stay the same because he failed to move. It stayed the same because he kept choosing to. Whether that distinction holds under scrutiny depends on what we can actually establish about his decisions - and we cannot establish much, because Howard did not explain himself.
What we can establish is narrower and more concrete. When the systems migration ran into the problem no one in the room had seen before, someone called Howard at 7pm. When a junior analyst drafted a proposal that was technically sound but politically unreadable, someone sent it to Howard first. When the team could not remember why a particular agreement was structured the way it was, Howard could cite the meeting where it was decided and name everyone in the room. These are not contested claims. They are the kind of evidence that accumulates over time and becomes difficult to explain away.
The harder question - what happens to that knowledge now, and whether it can be transferred - we have not answered and may not be able to. The onboarding documentation exists. Whether it captures what Howard carried is not yet established. That is not a failure of the effort; it may be a limit of the form.
What we cannot know is what the next crisis looks like when the person who would have known what to do is no longer on the call list.
Fourteen months of parallel-running two checkout systems while rebuilding from scratch is not the kind of project that generates a clean story. Let me try to say what the evidence actually supports.
What is established: the new system is live, and it held under peak load - the rollout did not fail, which was the outcome we least wanted and had most reason to fear after two slipped launch dates and at least two near-misses the team absorbed without the organization fully registering the cost. Priya’s call in the third quarter, to delay rather than ship a system she did not trust under concurrent write load, was the right call. We know that now because the alternative would likely have meant a corrupted order queue. We can say that with confidence.
What is not yet established: whether rebuilding the checkout actually fixes the abandonment problem. We have the new system in production. We do not yet have the evidence that it addresses the root cause. The usage data over the next two quarters will tell us whether the underlying friction was architectural or something else entirely. It would be premature to declare that problem solved.
What deserves to be said plainly: the team carried something most people did not see. Running dual systems for fourteen months, absorbing a twice-slipped launch without losing trust internally, navigating the near-miss in month eleven when Marcus rewrote the payment-event handler under deadline - this cost people time and reserves they will not get back. The work is done. That is what can be said, and it is enough.
The case for returning everyone to the office five days a week rests on claims that have not been established in any rigorous sense. Office-first advocates argue that presence builds trust and enables unplanned collision that produces ideas. Both may be true. Neither has been shown to be true in the specific situations where this policy would apply.
What we have observed - informally, in our own teams - is that some collaboration moments benefit from shared physical space, and others show no obvious difference. We cannot yet say which types of work belong in which category. The assertion that “being together more” produces better outcomes is plausible, but it has not been demonstrated for our context. If someone could show me which collaboration patterns actually require co-location and which do not, I would update my position.
The fully-remote position carries its own unverified assumptions. “Remote widens the talent pool” is a reasonable hypothesis. Whether it produces the specific talent this organization needs, and whether those hires integrate as effectively as local ones, is not something we have data on. That some individuals thrive remotely is established. That entire teams operate at the same quality without any shared context is a different claim, and it has not been shown.
The hybrid I am proposing does not require believing either set of unproven claims. It asks only that we protect the situations where presence has shown some advantage - cross-functional planning, new-member onboarding, periods where misalignment has compounded - and leave the rest flexible. If the anchor days turn out not to matter, the evidence will tell us. If full flexibility erodes something we cannot yet name, we will have the chance to notice. That is not a bold claim. It is the minimum defensible one.
Tidemark is a tool designed to pull scattered customer feedback into a single ranked, shareable roadmap. It launches next week.
The problem it addresses is real enough: small teams frequently manage customer input across a chat tool, a ticket tracker, a few shared documents, and whatever the last customer call left in someone’s notes app. Whether consolidating that input into one ranked list actually resolves the prioritization disagreements that drive teams toward spreadsheet chaos - that is a different and harder question.
What Tidemark claims to do is clear enough: ingest feedback from multiple sources, surface patterns, and produce an ordered list the team can share with stakeholders. The mechanism is plausible. Whether teams will maintain the discipline to feed it consistently - and whether the ranking algorithm’s implicit weighting will match any given team’s actual priorities - has not yet been established.
A few things to hold in mind before drawing conclusions. The absence of a better tool does not confirm that this one works. Teams that struggle to aggregate feedback often struggle because they disagree about what the feedback means, not merely where it lives. A shared document does not resolve that disagreement; it concentrates it. If Tidemark’s ranking reflects the team’s actual priorities, it will be useful. But whether those two things align depends on assumptions the tool cannot verify on the team’s behalf.
The launch next week will provide a clearer test. The question worth watching: Does the ranking surface surprises the team actually agrees with, or does it confirm what whoever controlled the weighting already believed? If the latter, the tool has automated a bias, not resolved one.
The year was hard. That much is not in dispute. What I am less certain about is the story I keep trying to build around it.
The project - 14 months of work that ended without the outcome I intended - gives me two competing accounts and I cannot rule out either. One says the timing was wrong, the audience wasn’t ready, the key people who might have carried it were looking elsewhere. The other says the work itself was not good enough, or that I had read the situation wrong from the start. I notice I prefer the first account. That preference is worth examining. I do not have enough data to weigh one over the other, and I am suspicious of myself on this point.
The relationship with Elise changed in ways I did not choose. I have a theory about why - actually I have several theories, and they contradict each other. The evidence I can access is a handful of conversations, a long silence, a few things said at a moment of strain. That is not enough to reconstruct causation from. It might have been something I did. It might have been something I had been, for a long time, without noticing. Absence of a clear explanation is not the same as the explanation being one thing rather than another.
What I am carrying forward: the questions more than the conclusions. If I found myself satisfied with a tidy account of this year, I would want to know what I had settled for.
Appears in diff-pairs
Section titled “Appears in diff-pairs”- skeptical vs candid (varies tone)
- skeptical vs candid (varies tone)
- skeptical vs candid (varies tone)