There are several paths to starting a career in software development, including the more non-traditional routes that are now more accessible than ever. Whether you're interested in front-end, back-end, or full-stack development, we offer more than 10,000 resources that can help you grow your current career or *develop* a new one.
Everybody Wants to Be a Dev!
Exploration vs Exploitation: Why It Matters and the Engineer’s Role
I've spent the better part of two years watching teams ship agentic AI systems, and a pattern keeps repeating. Two engineers read the same LangChain docs, attend the same conference talks, and build systems that look identical in a demo. Six months later, one system is handling thousands of requests a day with predictable behavior. The other gets quietly replaced by a simpler rules engine after it embarrassed someone in front of a customer. The gap between those two outcomes has almost nothing to do with model choice or framework familiarity. It comes down to a small set of skills that don't show up on most job postings for AI engineers, and that most online courses skip entirely. What Makes an Agentic AI System Different From a Chatbot Wrapper A chatbot wrapper takes input, sends it to a model, and returns the output. An agentic system makes decisions across multiple steps, calls tools, holds state, and sometimes calls itself. That difference sounds small written down. In practice, it changes everything about how the system fails. A wrapper that gives a bad answer wastes one turn. An agent that makes a bad decision at step two can compound that mistake across steps three through fifteen, calling the wrong API, writing bad data to a database, or looping on a task it can't complete. The failure modes are different in kind, not just in severity, and engineers who haven't built agentic systems before tend to debug them like they would debug a single bad response. They look at the final output instead of the decision trail that produced it. Skill One: Building Evaluation Before Building Features Most teams build the agent first and figure out how to test it later. The engineers who ship reliable systems do the reverse. Before writing the orchestration logic, they write a set of test cases the agent has to pass, with clear pass and fail criteria, and they run those cases against every change to the prompt, the tool definitions, or the model version. This sounds obvious stated plainly. It's rare in practice because agentic systems resist the testing patterns engineers already know. A unit test checks one function against one expected output. An agent's output depends on the conversation history, the tools available at that moment, and the specific phrasing of the user's request, so a single test case doesn't generalize the way a unit test does. Engineers who handle this well build small evaluation harnesses early, often before the agent does anything useful. They run twenty or thirty scenarios that represent the range of things the agent will see in production, including edge cases that look like they shouldn't happen. Then they track pass rate as a number they watch the same way they'd watch latency or error rate. When someone tweaks a system prompt to fix one issue, the harness catches the three other things that broke as a side effect. I've watched a team skip this step on a customer support agent, ship it, and discover three weeks later that a prompt change meant to improve tone had quietly disabled the agent's ability to escalate billing disputes to a human. Nobody caught it because nobody was running scenarios that exercised that path. A harness with even ten well-chosen test cases would have flagged it the same day. Skill Two: Treating Tool Definitions as an API Design Problem The tools an agent calls function as its only way of acting on the world, and most engineers write tool definitions the way they'd write internal function signatures: quick names, minimal descriptions, parameters that make sense to the person who wrote the code. That approach breaks down because the agent reads the tool description the same way it reads everything else, as natural language it has to interpret. A tool called search with the description "searches things" gives the model almost nothing to work with when it's deciding whether to call that tool or a different one, or what to pass as the query. Engineers who get this right write tool descriptions the way a technical writer would write public API documentation. They specify exactly when the tool should be used, what it returns, and what it doesn't do. They name parameters so the intent is obvious without a comment. A tool called search_customer_orders_by_email with a description stating it returns orders from the last 90 days and requires a verified email address gives the model far less room to misuse it than a generic search function does. This matters more as the number of available tools grows. An agent choosing between three tools can often guess right even with weak descriptions. An agent choosing between twenty tools, several of which sound similar, needs descriptions precise enough to disambiguate. Teams that scale past a handful of tools without revisiting this usually see a spike in wrong-tool-selected errors, and the fix is rarely a smarter model. It's better documentation. Skill Three: Designing for Partial Failure Traditional software either works or throws an exception. Agentic systems fail in a third way: the call succeeds, the response looks reasonable, and the content is wrong or incomplete. A tool call to fetch inventory data might return successfully while returning stale numbers. The model might decide a task is complete when it's only handled part of it. Engineers who've shipped production agents build explicit checkpoints into the flow where the system verifies its own progress against the actual goal, not just against whether the last API call returned a 200 status code. This might mean a verification step after a multi-stage task, where a separate prompt checks the agent's claimed output against the original request. It might mean structured outputs at each step that a deterministic function can validate, rather than trusting free text all the way through. The instinct to add more error handling here is correct, but the specific shape matters. Wrapping every tool call in a try-except block catches crashes. It doesn't catch an agent that confidently reports success on a task it didn't finish. That requires building verification logic that understands the task, not just the mechanics of the call. Skill Four: Knowing When Agentic Architecture Is the Wrong Choice The most consistent marker I've found for engineers who build agentic AI systems well is a willingness to argue against using one. 2026 has pushed agentic AI into the default answer for almost any automation problem, and that default is wrong often enough to matter. A task with a fixed sequence of steps and no real decision points doesn't need an agent reasoning through it each time. A deterministic pipeline runs faster, costs less, and fails in predictable ways that are easier to debug at 2 a.m. The engineers I'd trust with a production system are the ones who can look at a proposed agentic workflow and say plainly that a simpler architecture handles 90% of the cases just as well, reserving the agent for the genuine judgment calls. This isn't a popular position to take in planning meetings right now, with enterprise adoption of agentic AI accelerating across every sector and budget approval often tied to whether a project sounds sufficiently advanced. But the systems that hold up under real traffic tend to be the ones where someone pushed back on scope early, kept the agentic part narrow, and let boring code handle everything that didn't need a model making decisions. What This Looks Like Six Months In None of these four skills show up in a typical technical interview. They show up in incident reviews, in the difference between a system that degrades gracefully and one that fails in ways nobody anticipated, and in whether an engineer can explain why their agent made a specific decision three steps into a failed task. The teams shipping agentic systems that survive contact with real users aren't the ones with the most sophisticated prompts or the newest framework. They're the ones who treated evaluation as infrastructure, wrote tool descriptions like public documentation, built verification into the architecture instead of bolting it on after an incident, and stayed honest about when an agent was the wrong tool for the job. That combination is harder to hire for than "experience with LangChain" or "familiarity with RAG pipelines." It's also the actual difference between a demo and a system someone can depend on.
Leadership is challenging to develop in isolation. While you can practice programming, architecture, or databases independently, leadership relies on skills such as communication, influence, negotiation, feedback, conflict resolution, and decision-making, all of which require interaction with others. As leadership becomes more important for software engineers advancing in their careers, a key question arises: where can engineers practice these skills before becoming managers? Open source offers an ideal environment to develop both technical and leadership skills. Engineers tackle real technical challenges — such as coding, API design, architecture, testing, and documentation—while collaborating with individuals from diverse backgrounds, priorities, and perspectives. Although contributions often start with a pull request, advancing in the community requires explaining ideas, accepting feedback, building consensus, mentoring, and influencing technical direction. Open source is therefore more than a platform for technical growth; it serves as a practical setting for developing technical leadership. 1. Open Source as a Hard-Skill Accelerator For many software engineers, developing hard skills is a natural starting point. We are often eager to learn new languages, understand frameworks, enhance design skills, or explore different architectures. Open source offers a rich environment for this growth by exposing you to real software, real constraints, and ongoing evolution. Rather than working on isolated exercises, you can study and contribute to systems that have endured years or even decades of change. A key lesson is learning to manage software over the long term. Projects like Java, which have evolved for decades, reflect decisions about backward compatibility, modernization, deprecation, migration, performance, security, and ecosystem stability. This contrasts with greenfield applications, where ideas can be replaced freely. Mature open-source projects show that good engineering often means safely evolving an imperfect but widely used system, rather than aiming for perfect design. Open source provides practical experience with legacy modernization. You can observe how maintainers introduce new APIs without disrupting existing users, gradually remove obsolete abstractions, use tests to protect behavior during refactoring, and break down architectural changes into manageable steps. These challenges are common in enterprise environments but are difficult to replicate in personal projects. Another important area is documentation. In open source, documentation is not secondary to the code. API documentation, design discussions, migration guides, issue descriptions, proposals, release notes, and contribution guidelines are part of the engineering work itself. Writing clearly forces you to explain not only what the code does, but also why a decision exists and what trade-offs were considered. That ability becomes increasingly important as you move toward Staff Engineer or Architect responsibilities. Open source also offers opportunities to improve your coding and software design skills. You can study code written by engineers from diverse companies, countries, and technical backgrounds. This exposure is valuable because there is no single universal style of good software design. Projects optimize for different constraints, such as performance, compatibility, simplicity, extensibility, security, developer experience, or operational stability. Comparing these decisions helps you develop sound judgment rather than simply memorizing patterns. This is especially relevant in software architecture, where decisions are rarely clear-cut. Most architectural choices are shaped by context, constraints, history, and trade-offs. Open source allows you to observe these decisions openly, including API discussions, rejected proposals, compatibility concerns, implementation limitations, and competing approaches. You can see both the final architecture and the reasoning behind it. Open source offers a unique learning advantage: you can learn directly from the creators of the technologies you use. Instead of relying solely on tutorials or books, you can read their code, follow design discussions, review pull requests, and sometimes ask questions directly. Over time, you may even become one of the contributors shaping the project. Finally, understanding the internals of a framework, library, language, or specification can set you apart. Many engineers know how to use a technology, but few understand why it behaves as it does, its limitations, or its internal workings. Open source provides access to this deeper knowledge. For experienced software engineers, this understanding can make a significant difference when debugging complex issues, evaluating trade-offs, or making architectural decisions. 2. Open Source as a Soft-Skill Laboratory Many software engineers focused on technical expertise may overlook soft skills, assuming communication, persuasion, networking, and public speaking are primarily for managers. However, advancing in a technical career requires these abilities. Software is built collaboratively, key decisions are made through discussion, and achieving greater impact depends on others understanding, trusting, and supporting your ideas. Open source offers a practical environment to develop these skills, as the outcomes are tangible. You propose changes, defend technical decisions, receive feedback, collaborate with unfamiliar colleagues, and work to make your ideas clear and accepted by others. Learn to Communicate Through Writing A significant amount of software engineering leadership happens in writing. Issues, pull requests, design proposals, mailing lists, documentation, specifications, and code reviews all require you to organize your thoughts before requesting action. Open source provides frequent opportunities to practice this skill. This skill extends beyond open source. For example, the value of an Architecture Decision Record relies on your ability to describe context, explain alternatives, clarify trade-offs, and ensure the decision is understandable to future readers. Writing is not just documentation; it transforms technical reasoning into content that can be shared, challenged, and reused. Learn to Explain and Sell Technical Ideas Technical leadership also requires speaking. You may need to defend architectural decisions, explain preferred designs, challenge existing approaches, or persuade multiple teams to adopt new directions. Having an idea is only the first step; you must also make it understandable to those without your context. Open-source communities offer many opportunities to practice this: community calls, meetups, user groups, podcasts, workshops, and conferences. Preparing a presentation requires you to organize complex information, remove unnecessary details, build a clear narrative, and explain your reasoning so others can follow. That ability is crucial for any senior software engineer. Communicate Across Languages and Cultures Open source is global. If English is not your first language, as it is not mine, participating in international communities offers ongoing opportunities to improve. You regularly write issues, join discussions, review proposals, attend meetings, and present ideas in English. But the learning goes beyond vocabulary or grammar. You also learn how people from different cultures communicate, disagree, provide feedback, and make decisions. What seems normal in one culture may appear aggressive or ambiguous in another. For engineers in global organizations, effective cross-cultural communication can be as important as learning a new technical framework. Build Relationships and Reputation Open source can expand your network organically. You do not meet people simply to “network.” Instead, others recognize you through your consistent work, contributions, reviews, and participation in discussions. Over time, people learn your expertise and know what they can rely on you for. This is valuable because reputation extends beyond organizational boundaries. By sharing knowledge through technical decisions, pull requests, articles, documentation, or presentations, you can help people outside your company see your approach. External credibility can also strengthen your reputation within your organization. However, building reputation is a long-term investment. A few pull requests or a single conference talk will not transform your career. Reputation develops over months and years through consistent contributions. Learn to Manage Your Time and Context Open source also helps develop an underrated leadership skill: managing your attention. Most engineers contribute to open source while managing full-time jobs and other responsibilities. This requires deciding what deserves your time, breaking large initiatives into smaller tasks, prioritizing contributions, and switching contexts efficiently. These skills become increasingly important as your career progresses. Staff Engineers or Architects rarely focus on a single task. They often move between architecture discussions, code reviews, mentoring, incidents, multiple teams, and long-term initiatives within the same week. Maintaining focus while working across multiple contexts becomes essential. Build Discipline Through Consistency Open source also fosters discipline. While large contributions are visible, sustainable open-source involvement is built through smaller actions such as reviewing issues, improving documentation, answering questions, writing tests, fixing bugs, or joining design discussions. Success rarely comes from a single heroic contribution. It is consistency. Consistently doing small, meaningful work leads to long-term growth. You gain a deeper understanding of the project, earn recognition, take on more responsibility, and may eventually help shape the technology’s direction. The same principle applies to leadership. Leadership develops through repeated opportunities to communicate, influence, help others, make decisions, and earn trust, not simply by receiving a title. Open source simply gives you many more opportunities to practice. Conclusion A strong software engineer must develop both technical expertise and leadership skills to become a well-rounded professional. Excelling at coding, system design, or architecture is not enough if you cannot navigate challenging discussions, communicate with stakeholders, build trust, and clearly explain your ideas. Good technical ideas often fail when they are not understood, trusted, or convincingly presented. The reverse is equally risky. Strong communication and influence, without sufficient technical foundation, can lead teams astray. Leadership without technical judgment may result in persuasive presentations built on weak decisions. Conversely, technical depth without leadership can keep valuable ideas from being realized. High-impact engineering demands both skill sets. This balance is essential for those pursuing roles such as Software Architect, Staff Engineer, Principal Engineer, or technology executive. Complete knowledge is not expected. The key skill is the ability to shift between strategic discussions with C-level leaders and technical conversations with engineers to understand implementation details and design trade-offs. Open source offers valuable opportunities to develop both technical and leadership abilities, helping engineers grow as technologists and leaders.
Many assume that leadership in software engineering starts only when you stop coding and become a manager. I once shared this belief, thinking technology would be simpler than working with people. That was an early misconception. While it is possible to focus your career on code, architecture, databases, and other technical areas, the challenges that shape your impact become less technical over time. Even the best architectural decision has little value if others do not trust, understand, or support it. This does not mean every experienced software engineer should become a manager. Leadership is equally important on the technical track. Senior individual contributors, such as staff engineers, architects, and principal engineers, are expected to influence decisions beyond their own code. As Will Larson discusses in Staff Engineer, advancing beyond senior engineering focuses on technical leadership rather than people management. To increase your technical impact, others must listen to your ideas, trust your judgment, include you in key discussions, and act on your recommendations. You may choose not to manage people, but avoiding leadership will eventually limit your growth as a software engineer. What Do We Mean by Leadership? Leadership predates corporations, job titles, and management frameworks. For example, the Roman military’s success relied not only on superior weapons or armor but also on effective organization. Legions were divided into smaller units, each with defined responsibilities and led by centurions. Leadership was distributed throughout the ranks, enabling coordinated efforts toward larger objectives. A similar concept appears in the term architect, commonly used by software engineers. Derived from the Greek arkhitekton — arkhi meaning chief and tekton meaning builder — an architect was the master builder, responsible for both understanding the craft and directing others. This role closely resembles that of an effective software architect today. Another example comes from the nineteenth-century Prussian military, which developed its General Staff as a professional body focused on planning, coordination, and operational readiness. This approach recognized that complex organizations require skilled individuals to address critical challenges without making each one the commander. The model became influential and was adopted by other militaries. Software engineering faced a similar challenge: how can experienced engineers expand their organizational impact without moving into people management? An early solution appeared in the British Royal Navy, where managing large fleets required separating command authority from technical expertise. Naval operations relied on both captains and skilled officers responsible for navigation, planning, logistics, and coordination. This staff function supported fleet-level decision-making without direct command and influenced how organizations approach distributed expertise and coordination. Modern software organizations independently adopted a similar approach. Titles such as Staff Engineer, Principal Engineer, and Distinguished Engineer now represent technical leadership roles. Will Larson highlights this distinction in his 2021 book, Staff Engineer: Leadership Beyond the Management Track, which explores Staff-plus engineering as leadership outside the traditional management ladder. This distinction is essential: management is a role, while leadership is an activity. Managers have formal responsibilities for people, performance, hiring, priorities, and processes. Technical leaders may lack formal authority, but their influence comes from expertise, judgment, communication, trust, and guiding better technical decisions. In software engineering, leadership does not require direct reports; it often means being the trusted engineer who provides direction in complex situations. Why Should a Software Engineer Care About Leadership? While understanding leadership is valuable, it is even more important to consider why a Software Engineer who does not plan to become a manager should invest time in developing these skills. As your career advances, your impact depends not only on your technical skills but also on your ability to influence decisions, collaborate effectively, and guide the organization toward better technical outcomes. 1. Software Development Is About People Software development is inherently a social activity. Software is built collaboratively with engineers, product managers, designers, architects, clients, managers, and other stakeholders. Even highly technical decisions must be explained, discussed, challenged, negotiated, or approved by others. A database migration may seem purely technical until it impacts another team. An architectural decision becomes a communication challenge when multiple teams must adopt it. Even an elegant solution can fail if it does not address the client’s real needs. While much of your day may involve working with machines, software exists for people, is created by people, and ultimately serves people. Choosing not to pursue management does not eliminate the human aspect of software engineering. 2. Technical Expertise Needs Trust, Influence, and Access Technical correctness alone is not sufficient. You might understand why a particular architecture will not scale, recognize that a technology introduces unnecessary complexity, identify an important security risk, or propose a significantly better design. However, your expertise has limited impact if others do not listen. Technical expertise leads to organizational impact only when you can influence the organization’s actions. Influence rarely stems from expertise alone. People must trust your judgment, see that you understand the context, listen to opposing views, explain trade-offs clearly, and adapt your position when evidence changes. This is also why relationships matter in a technical career. As your responsibilities increase, many key decisions occur outside the codebase, such as during architecture reviews, design discussions, planning sessions, incident reviews, roadmap meetings, and cross-team or stakeholder meetings. If you want to influence those decisions, you need to be part of those conversations. Leadership helps you build the credibility and trust needed to participate in these discussions and ensures your voice is heard. A strong technical leader does more than provide correct answers. They create conditions where good technical decisions can be understood, challenged, accepted, and implemented. 3. If You Do Not Lead, Someone Else Will Make the Decision When experienced engineers avoid leadership, an uncomfortable consequence arises. The decisions do not disappear. Someone else will make them. And that person may have considerably less technical understanding of the consequences. This can lead organizations to measure engineering productivity using questionable proxies such as lines of code, number of commits, tickets closed, or tokens consumed by AI tools. When engineers encounter such decisions, their natural reaction is often: “Who thought this was a good idea?” A better question might occasionally be: “Which experienced engineers were involved when this decision was made?” Leadership ensures that technical knowledge is represented in decisions affecting engineering. You do not need to control every decision, but you should be willing to participate in the important ones. 4. Leadership Multiplies Your Technical Impact There is a natural limit to how much software one person can build individually. Even exceptional engineers have limited time each day. Leadership enables your expertise to extend beyond those limits. By helping others make better design decisions, establishing reusable architectural approaches, mentoring engineers, improving practices, or preventing costly mistakes, your impact exceeds your individual contributions. This marks an important transition in senior technical careers. Early in your career, your value is largely based on your individual contributions. Later, your value increasingly comes from enabling other engineers and teams to succeed. Your code remains important, but it is no longer the sole measure of your contribution. 5. Leadership Becomes Part of Technical Career Progression Leadership becomes increasingly integral to career progression beyond the Senior Software Engineer role. Staff Engineers, Principal Engineers, Distinguished Engineers, and Software Architects may remain Individual Contributors, yet their responsibilities typically extend beyond implementing individual features. They are expected to provide technical direction, navigate ambiguity, resolve difficult trade-offs, connect teams, mentor engineers, challenge assumptions, and influence decisions whose consequences may extend across an organization. None of those responsibilities inherently requires becoming a people manager. But almost all of them require leadership. Treating leadership as exclusive to management can eventually limit the career growth of experienced Software Engineers. You can choose not to manage people. You can choose to remain deeply technical. As your scope and impact grow, leadership increasingly becomes integral to technical work. Conclusion Leadership in software engineering does not mean leaving the technical path or becoming a manager. It means understanding that software is built by people, and that technical expertise must be paired with trust and influence to drive change. Key decisions need experienced engineers involved, and leadership enables your knowledge to reach beyond your own code. As you advance to roles like Staff Engineer, Principal Engineer, or Software Architect, leadership becomes essential for greater impact. You can remain an Individual Contributor and stay deeply technical, but to increase your influence, you must also guide, communicate, build trust, and help others make better decisions.
Almost every backend eventually needs to run code on a schedule. Send the invoice at midnight. Retry the failed payment in five minutes. Generate the weekly report every Monday at 7 AM. Clean up expired sessions every hour. On one server, this is easy. You write a cron line and move on. The trouble starts when one server becomes ten. Now the same cron line lives on every box, so the invoice job fires ten times instead of once. Move the cron to a single “scheduler” box, and that box becomes a single point of failure. Every time you deploy new code, that process restarts, and if it crashes or the host dies, there is no second node to cover for it. Any job due during that downtime window silently never fires. A distributed job scheduler solves this. It runs jobs reliably across a fleet of machines, fires each job once even when nodes crash, and keeps working when parts of the system fail. This post walks through how to design one, the trade-offs at each step, and the mistakes that bite teams in production. What the Scheduler Has to Do Before drawing boxes, it helps to pin down the requirements. They split into two groups. Functional requirements: Run a job once at a specific time (a one-time job).Run a job on a repeating schedule, usually defined with cron (a recurring job).Support job dependencies, where job B runs only after job A succeeds.Retry a job automatically when it fails.Respect priority, so urgent jobs run before bulk jobs.Cancel or pause a job that is scheduled or already running. Non-functional requirements: Durability. Once the system accepts a job, it must not lose it, even if a node dies one second later.At-least-once execution. Every due job runs at least one time.Scale. The design should handle millions of jobs per day across many workers.Fault tolerance. A crashed worker must not block other jobs, and its work should be picked up by someone else. One requirement is worth calling out early. People often ask for “exactly-once” execution. In a distributed system, you cannot truly get it. What you can build is at-least-once delivery plus idempotent jobs, which together behave like exactly-once from the outside. More on that later. The Core Architecture The single most important idea in this design is to separate deciding when a job runs from actually running it. These are two different problems with different scaling needs, so they become two different components. A clean design has four parts: A scheduler that watches the clock and decides which jobs are due.A queue that holds ready-to-run jobs and hands them out.A pool of stateless workers that pull jobs and execute them.A datastore that holds job definitions and execution history, and acts as the source of truth. Why decouple the queue from the workers at all? Because load is bursty. At midnight, a thousand daily jobs may become due at the same second. If the scheduler called workers directly, that spike would hit them all at once. The queue absorbs the spike and lets workers drain it at a steady rate. It also lets you scale workers up and down without touching the scheduler. This is the same reason queues show up across system design, which I covered in detail in Role of Queues in System Design. Modeling Jobs in the Database The datastore is the source of truth, so the schema matters. A common approach uses two tables. One holds the recurring definition, the other holds individual runs. SQL CREATE TABLE jobs ( id BIGINT PRIMARY KEY, name TEXT NOT NULL, cron TEXT, -- null for one-time jobs payload JSONB, next_run_at TIMESTAMPTZ, -- when this job is next due enabled BOOLEAN DEFAULT TRUE ); CREATE TABLE job_runs ( id BIGINT PRIMARY KEY, -- unique id per run job_id BIGINT REFERENCES jobs(id), status TEXT NOT NULL, -- PENDING, RUNNING, SUCCEEDED, FAILED, DEAD attempt INT NOT NULL DEFAULT 1, scheduled_at TIMESTAMPTZ, started_at TIMESTAMPTZ, lease_until TIMESTAMPTZ ); CREATE INDEX idx_jobs_due ON jobs (next_run_at) WHERE enabled = TRUE; The partial index on next_run_at is the workhorse. The scheduler asks “which jobs are due now” many times per second, and this index keeps that query fast even with millions of rows. Each run moves through a small set of states. Drawing the state machine makes the retry and failure logic obvious. Defining Schedules With Cron Recurring jobs need a way to express “every day at 2:30 AM” or “every 15 minutes.” Cron is still the standard. A classic cron expression has five fields: Plain Text minute hour day-of-month month day-of-week 30 2 * * * -> 2:30 AM every day The Java world often uses Quartz cron, which adds a seconds field at the front and a year field at the end, giving six or seven fields. The two formats look similar but are not interchangeable, and mixing them up is a frequent source of jobs that never fire. The scheduler stores the cron string and computes a concrete next_run_at timestamp from it. After a run is enqueued, it computes the next one. This raises a real question: what happens if the scheduler was down for an hour and three runs were missed? This is the misfire problem. You generally pick one of two policies: Catch up. Run every missed occurrence in order. Correct for billing, expensive for everything else.Skip. Run only the next future occurrence and forget the missed ones. Right for jobs like cache refreshes where stale runs add no value. Make this an explicit setting per job. Teams that leave it implicit get surprised after the first outage. Picking Which Jobs to Run The scheduler needs to find due jobs and hand them off. There are three common ways to find them. Polling. Every second, query the database for jobs where next_run_at <= now(). Simple and reliable. The partial index keeps it cheap. The cost is a small delay, up to your poll interval.Timer wheel. Keep upcoming jobs in an in-memory structure sorted by time. Very precise and great for short delays, but you have to rebuild it from the database after a restart.Push. An external timing service fires an event when a job is due. Real-time, but now you depend on another moving part. For most systems, polling with a one-second interval is the right default. It is boring, and boring is good for a component you are trusting with billing runs. The harder problem is concurrency. If you run several scheduler instances for availability, they will all poll the same table at the same time. Without care, two of them pick the same job, and it runs twice. The clean fix in PostgreSQL is row locking with SKIP LOCKED: SQL SELECT id FROM jobs WHERE enabled = TRUE AND next_run_at <= now() ORDER BY next_run_at LIMIT 100 FOR UPDATE SKIP LOCKED; FOR UPDATE locks the rows this instance selects. SKIP LOCKED tells other instances to ignore locked rows and grab the next free ones instead. Many schedulers can now poll in parallel, each claiming a different batch, with no coordination service and no duplicate pickups. Airflow uses exactly this approach instead of a heavier consensus protocol, which is a good reminder that the simplest mechanism that meets the requirement usually wins. Why Exactly-Once Is a Myth Here is the scenario that breaks naive designs. A worker pulls a job, runs it successfully, and then crashes before it can tell the system “done.” The system still thinks the job is running. The lease expires, another worker picks it up, and the job runs a second time. You charged the card twice. You cannot delete this scenario. Networks drop messages and processes die at the worst moment. So you stop chasing exactly-once delivery and instead make the work safe to repeat. That means two things working together: At-least-once delivery. The system guarantees a due job runs at least one time, accepting that it may occasionally run more than once.Idempotent jobs. Running the same job twice has the same effect as running it once. The standard trick is an idempotency key built from stable identifiers, for example {job_id, run_id, attempt}, or a key tied to the business action like invoice_2026_06_charge. The worker records that key before committing side effects. If the same key shows up again, the worker sees the work is already done and acknowledges without repeating it. This is why each run gets its own unique id. A time-ordered id such as a Snowflake id or a ULID works well, because it is unique across the whole fleet without coordination and it sorts by creation time, which keeps the job_runs table naturally ordered. I explained the structure of these ids in How Snowflake IDs Work, and the deduplication pattern itself in Idempotent Receiver Pattern. There is one more subtle gap. The worker has to update the database and publish to the queue, and those are two systems. If it writes to the database and then dies before publishing, the job is lost. The transactional outbox pattern closes this gap by writing the job and an outbox row in one local transaction, then publishing from the outbox separately. I covered that in The Transactional Outbox Pattern. Coordinating at Scale A single scheduler instance has a throughput ceiling. Past a certain number of jobs per second, one process polling one database cannot keep up. There are two ways to grow. The first is leader election. You run several scheduler instances, but only one is active at a time. The others stand by and take over if the leader dies. A coordination service like etcd or ZooKeeper holds the leadership lock. This is simple to reason about, but the single active leader is still a throughput bottleneck. The second is sharding. You split the job space across many active schedulers. A simple scheme hashes the job id into one of N partitions, and each scheduler owns a set of partitions. Every job has exactly one owner, so there are no duplicate pickups, and throughput grows by adding schedulers. Consistent hashing makes it cheaper to add or remove schedulers without reshuffling everything. Sharding has one sharp edge. During a handover, while leases for a partition are changing hands, two schedulers can briefly believe they own the same partition. This is split brain. You do not try to make it impossible, because that is expensive. Instead, you let the worker-side idempotency check be the final safety net. If both schedulers enqueue the same run, the idempotency key means it still executes once. Google’s cron service takes a stricter route for its most sensitive launches. It writes the launch record to a quorum using Paxos before the job actually starts, so a failover cannot lose or double-fire it. For most teams, leases plus idempotency are enough, and full consensus is overkill. Detecting Failures and Recovering Workers crash. The scheduler has to notice and reassign their work, without stealing jobs from workers that are simply slow. The mechanism is a lease with a heartbeat. When a worker claims a run, it sets lease_until to a short time in the future, say 30 seconds. While the job runs, the worker periodically extends the lease. If the worker dies, it stops extending, the lease expires, and a recovery sweep moves the run back to PENDING so another worker can take it. SQL -- recovery sweep: reclaim runs whose lease has expired UPDATE job_runs SET status = 'PENDING' WHERE status = 'RUNNING' AND lease_until < now(); Two details make this robust. First, the lease timeout must be comfortably longer than a normal heartbeat interval, or a brief pause will cause a healthy job to be wrongly reclaimed. Second, you need protection against a zombie worker, one that froze on a long garbage collection pause, lost its lease, and then woke up and tried to finish writing results. A fencing token solves this. The reclaimed run gets a higher token, and the datastore rejects any write carrying an older token. I went deeper on time-bound ownership and fencing in The Lease Pattern in Distributed Systems. Retries Done Right A failed job should usually be retried, but retrying badly makes outages worse. If a downstream service is struggling and every failed job retries immediately, you pile on more load at the exact moment it can least handle it. The fix is exponential backoff with jitter. Each retry waits longer than the last, and a random jitter spreads the retries out so they do not all fire at the same instant. Plain Text attempt 1 fails -> wait ~1s attempt 2 fails -> wait ~2s attempt 3 fails -> wait ~4s attempt 4 fails -> wait ~8s (each wait randomized by +/- a few hundred ms) After a fixed number of attempts, stop. A job that keeps failing should not retry forever. Move it to a dead letter queue, a separate place for runs that exhausted their retries, and alert a human. The dead letter queue keeps a poisoned job from clogging the pipeline while preserving it for investigation. Operating the Thing A scheduler is infrastructure other teams depend on, so it has to be observable and controllable. For observability, track the metrics that tell you the system is healthy: Queue depth. A queue that keeps growing means workers cannot keep up.Scheduling lag, the gap between when a job was due and when it actually started.Run outcomes per minute, split by succeeded, failed, and dead.Lease reclaims, which spike when workers are crashing. For control, give operators real knobs. They should be able to pause a queue, drain a worker before a deploy so it finishes current jobs and takes no new ones, and replay a dead-lettered job after fixing the cause. Building these in from the start saves a lot of pain during the first incident. How Real Systems Approach This None of this is theoretical. The same building blocks show up across well-known tools, each making a different trade-off. Quartz. A mature Java scheduler. Multiple instances coordinate through a shared database using row locks, the same idea as the SKIP LOCKED approach above.Airflow. Orchestrates dependency graphs of tasks. Its scheduler uses database locks rather than a consensus protocol, favoring operational simplicity.Temporal. Models workflows as code and replays an append-only event history to recover state after a crash, which sidesteps a whole class of mid-task failure bugs.Celery. A popular task queue in Python, with a beat component that handles periodic scheduling.Kubernetes CronJobs. Run containerized jobs on a cron schedule inside a cluster, with configurable policies for missed runs and concurrency. See the Kubernetes CronJob docs.Google distributed cron. Writes launch state to a Paxos quorum before launching, so a leader failover never loses or doubles a run. The pattern across all of them is consistent. Decouple scheduling from execution, lean on the database or a quorum for coordination, accept at-least-once and make jobs idempotent, and design for failure as the normal case. Takeaways If you remember five things from this, make it these. Separate the decision of when a job runs from the work of running it. They scale differently.Do not chase exactly-once. Build at-least-once delivery and make every job idempotent.Use the database as a coordination primitive. SELECT ... FOR UPDATE SKIP LOCKED lets many schedulers poll safely.Use leases with heartbeats and fencing tokens to detect dead workers and reclaim their runs without double execution.Retry with exponential backoff and jitter, cap the attempts, and send the rest to a dead letter queue. A good scheduler is not clever. It is careful. It assumes nodes will die, messages will duplicate, and clocks will drift, and it keeps running anyway.
Artificial intelligence is rapidly transforming software testing by enabling QA engineers to generate test cases and test plans, automate browser interactions, analyze and debug failures, and execute complex testing workflows using simple natural-language prompts. While cloud-based AI assistants offer impressive capabilities, they often require subscriptions and sharing potentially sensitive application data with third-party services. Running an AI-powered testing assistant locally addresses these concerns by providing better privacy, lower operating costs, and complete control over the testing environment. In this tutorial, we’ll learn how to build our own local AI QA engineer using Docker, Ollama, Qwen3:8b, LibreChat, and Playwright MCP. It will allow us to perform browser automation and interact with web applications using natural language, all without relying on cloud-based AI services. Understanding the Architecture Every interaction begins with the user. For example, a user enters a prompt in LibreChat, such as “Open the Playwright website and click the ‘Get Started’ button.” LibreChat serves as the conversational interface through which users interact with the AI assistant. Rather than processing the request itself, it forwards the prompt to a locally hosted large language model, Qwen3:8b, running via Ollama. After receiving the prompt, Qwen3:8b interprets the user’s intent and generates a step-by-step execution plan. Instead of interacting with the browser directly, the model determines which tools are required and communicates those instructions using the Model Context Protocol (MCP). These MCP requests are handled by the Playwright MCP Server, which acts as the bridge between the language model and the browser. It translates the AI-generated instructions into executable Playwright commands. The Playwright MCP Server then launches a Chrome browser and performs the requested actions. Depending on the prompt, it can navigate to websites, click buttons, complete forms, extract text from web pages, capture screenshots, and execute a wide range of browser automation tasks. Once the browser completes the requested operations, the execution results are returned to Qwen3:8b. The language model analyzes the browser output and transforms the technical details into a clear, human-readable response. LibreChat then presents this response to the user. Instead of displaying raw Playwright logs, it provides a concise summary such as: “Navigation completed successfully. The Playwright website was opened, and the Get Started button was clicked successfully.” This architecture enables browser automation through natural language while ensuring that every component runs locally. As a result, we benefit from enhanced privacy, greater security, and complete control over the entire AI-powered automation workflow. Prerequisites Before getting started, ensure that the following software is installed on your machine: DockerNode.js 20 or higher versionGitOllama We’ll use Docker Desktop to run LibreChat, Node.js to install and run the Playwright MCP Server, Git to clone the required repositories, and Ollama to download and serve the local large language model. Having these tools installed beforehand will make the setup process smooth and straightforward. System Requirements Running a local AI-powered browser automation stack requires a reasonably capable machine. A system with 16 GB of RAM or more is recommended to run Docker containers and the language model efficiently. We’ll also need 20–25 GB of available disk space, preferably on an SSD, to accommodate Docker images and downloaded models. While a dedicated GPU can significantly improve model inference speed, it is entirely optional, and the setup works well on modern CPUs. For this tutorial, I’m using the following configuration: Operating system: macOS (M2 Pro)Memory: 16 GB RAM We can have the same setup on Windows and Linux, with only minor platform-specific differences in the installation steps. Setting Up the Environment for the Local AI QA Engineer Docker, Node.js, and Git are widely used development tools, and detailed installation guides for each are readily available online. Installing Ollama To install Ollama, either download the installer from the official website or use the installation command provided for your operating system. For macOS, it can also be installed using the following Homebrew command: Plain Text brew install ollama Once the installation is complete, it can be verified by running the following command in the terminal: Plain Text ollama --version Installing Qwen3:8b Qwen3:8b is chosen for this setup because it offers a strong balance of reasoning, code generation, and performance, making it ideal for Playwright TypeScript test generation, AI agents, MCP integration, and modern QA automation workflows while running efficiently on a local machine. However, other higher models can also be chosen if you know a better one. Another factor in choosing this model was the available system memory. Since my machine has 16 GB of RAM, some memory also needs to be reserved for other tools used in this setup, such as Docker, LibreChat, and Playwright. We need to start Ollama first by running the following command from the terminal. (It should be kept running in the background): Plain Text ollama serve Open a new terminal and run the following command to pull the Qwen3:8b model: Plain Text ollama pull qwen3:8b It should take some time to complete the pull, as the model is around 5.2GB. Once the download completes, we can check the model by running the command: Plain Text ollama list It should list the model downloaded. Next, we can quickly verify by running the model using the command: Plain Text ollama run qwen3:8b Once the model starts, it will prompt you to enter a query. To verify that everything is working correctly, try a simple prompt such as “What is 2 + 2?”. Observe how the model processes the request and generates its response. If the setup is successful, it should return the correct answer, 4, confirming that the model has been downloaded, installed, and is functioning properly. To stop the model, type “/bye” in the prompt, and it should exit. Qwen3:8b provides a good balance between performance and resource usage, making it a suitable choice for this hardware configuration. If more RAM is available, you can opt for larger LLMs that offer stronger reasoning and coding capabilities. Installing LibreChat With Docker LibreChat is an open-source AI platform that provides a unified and customizable interface for interacting with multiple AI models. It enables us to manage all our AI conversations from a single application while supporting features such as AI agents, Model Context Protocol (MCP) servers, custom tools, and integrations with both local and cloud-based LLMs. LibreChat acts as the front-end chat interface that communicates with the locally running Qwen3:8b model through Ollama. It allows us to execute AI-powered browser automation workflows entirely on our local machine. Follow the steps below to install LibreChat: Step 1: Clone the LibreChat GitHub Repository The repository can be cloned by running the following command: Plain Text git clone https://github.com/danny-avila/LibreChat After cloning the repository, navigate to the LibreChat folder, copy the .env.example file, and create a new .env file from it. Plain Text cd LibreChat cp .env.example .env Let's keep the .env file as it is, using the default values. Step 2: Connect Ollama to LibreChat Ollama can be connected to LibreChat by updating its configuration in the “librechat.yaml” file. The example file is already available in the cloned repo. Run the following command to copy librechat.example.yaml and create librechat.yaml. Plain Text cp librechat.example.yaml librechat.yaml Update the following configuration in the file to connect Ollama to LibreChat: YAML endpoints: custom: - name: "Ollama" apiKey: "ollama" baseURL: "http://host.docker.internal:11434/v1" models: default: - "qwen3:8b" fetch: true titleConvo: true titleModel: "current_model" summarize: false summaryModel: "current_model" modelDisplayLabel: "Ollama" Make sure that this configuration is added to the “custom” block, which falls under the “endpoints” block. This configuration adds Ollama as a custom AI endpoint in LibreChat. The baseURL tells LibreChat where to connect to the Ollama API, while the default model specifies that Qwen3:8b should be used by default. Since LibreChat is running inside a Docker container while Ollama is running directly on the host machine, we use http://host.docker.internal:11434/v1 instead of localhost. The special hostname host.docker.internal allows the Docker container to access services running on the host system, enabling LibreChat to connect to the locally running Qwen3:8b model through Ollama. Setting fetch: true allows LibreChat to automatically detect and display all models available in Ollama. The remaining options configure the user interface by generating conversation titles using the current model, disabling conversation summarization, and displaying the endpoint with the label Ollama in the LibreChat interface. Step 3: Mount the Configuration in the docker-compose-override.yml The docker-compose-override.yml can be copied and created in the same way as we did “librechat.example.yaml”. Plain Text cp docker-compose.override.yml.example docker-compose.override.yml The following block should be updated in the docker-compose.override.yml file. YAML services: api: volumes: - ./librechat.yaml:/app/librechat.yaml This file mounts the custom “librechat.yaml” configuration file into the LibreChat container. By mapping ./librechat.yaml to /app/librechat.yaml, Docker ensures that LibreChat uses the custom configuration each time the container starts. This approach allows us to modify settings, such as custom endpoints and AI models, without rebuilding the Docker image. Step 4: Start the LibreChat Application Using Docker Compose The LibreChat application can be started using the following command: Plain Text docker compose up -d It will take some time for the Docker images to download, and containers will start. Run the following command from the terminal to check the Container status: Plain Text docker ps -a This command displays the status of all Docker containers. If any container is unhealthy or encounters an issue, its status will be clearly indicated in the output. In case any container is unhealthy or encounters an issue, the following command can be run to check its logs: Plain Text docker logs
Landing a data engineering role means clearing a gauntlet that no other software discipline has to face all at once: airtight SQL, production-grade Python, data modeling instincts, distributed-compute fluency (Spark, warehouses, ETL), and system design that has to survive real data volume. Generic coding prep barely scratches the surface, and "just grind LeetCode" advice falls apart the moment an interviewer asks you to model a slowly changing dimension or reason about a skewed join. So we did the work. We evaluated the resources data engineers actually use, judged on five things that matter: relevance to the DE interview loop, depth of practice, realism of the questions, feedback quality, and price. Below is the ranked list. A quick note on methodology: this ranking favors resources that target the data engineering loop specifically, not generic algorithm grinding. That bias is intentional, and it is why the order may surprise you. 1. DataDriven.io Most "interview prep" platforms were built for generic SWE roles and bolt on a SQL section as an afterthought. This one was built from the ground up for the data engineering loop. The catchphrase you will hear repeated in DE communities is that DataDriven.io is LeetCode for data engineers, and it fits: instead of inverting binary trees, you are writing window functions against realistic schemas, designing star schemas, debugging an ETL transform, and reasoning about partitioning, all in an in-browser SQL and Python sandbox that runs your query against real data and tells you exactly where it broke. It is also the rare place where the whole product is built for the job rather than adjacent to it, which is why datadriven.io is great for data engineer interview prep specifically: SQL practice that ramps to multi-CTE analytics, a deep set of Python practice problems, plus data modeling, dimensional modeling, PySpark, and system-design tracks, with execution-based feedback and a difficulty curve that reaches the staff-level questions that actually separate offers from rejections. Verdict: The most targeted, realistic data engineering interview practice available today. Earns the top spot. 2. "Cracking the Coding Interview" (the book, by Gayle Laakmann McDowell) A deserved classic, and intentionally a book rather than a website. CTCI is still the best single artifact for understanding how technical interviews are actually structured: how the conversation flows, how to think out loud so the interviewer can follow your reasoning, how to recover when you get stuck, and how to handle the behavioral and negotiation segments that strong candidates routinely fumble. Most people lose offers not because they could not solve the problem but because they could not show their work, and this book is the canonical fix for that. Where it falls short for our purposes is scope. It will not teach you windowed SQL, slowly changing dimensions, or how to design a lakehouse, and its algorithm focus skews toward generalist software roles rather than the data engineering loop. The data structures and big-O chapters are still worth a pass because algorithm screens do show up, but treat them as a refresher, not your main event. Read CTCI once early in your prep to fix your interview mechanics, internalize the communication patterns, then spend the rest of your time on hands-on, domain-specific platforms. Verdict: Essential reading for interview mechanics; not a substitute for domain practice. 3. "Designing Data-Intensive Applications" (the book, by Martin Kleppmann) If CTCI teaches you how to interview, "DDIA" teaches you what a data engineer is actually supposed to know. Replication, partitioning, consistency models, batch versus stream processing, storage engine internals, the failure modes of distributed systems: this is the conceptual backbone of nearly every data engineering system design round. When an interviewer asks why you would choose a log-structured merge tree over a B-tree, or how you would keep two datastores in sync without losing events, the answers live in these pages. It is dense, and it is emphatically not an interview drill book. You will not find practice questions, and you cannot cram it the night before. What it gives you instead is judgment: the candidate who has internalized DDIA answers "how would you design this pipeline" with the calm of someone who has already thought through the tradeoffs, names the failure cases before being prompted, and explains why a choice holds up under real data volume. Read it slowly over weeks, ideally early in your prep, and pair it with a hands-on platform so the concepts attach to actual queries and schemas rather than floating as theory. Verdict: The definitive conceptual reference. Read it slowly, alongside real practice. 4. LeetCode The default destination, and it earns its spot for one practical reason: the Database problem set is sizable, the algorithm catalog is enormous, and the platform's brand means a large share of companies still pull their initial coding screen straight from it. If your target company is known to run a generic algorithm round before the data-specific rounds, you need exposure here, and the sheer volume of problems plus community discussion means you will rarely be surprised by a pattern you have never seen. The catch for data engineers is that LeetCode was built for the algorithm interview, not the DE loop. Its SQL section is genuinely solid but secondary; the questions are puzzle-shaped rather than drawn from real schemas, and you will not find data modeling, ETL design, dimensional modeling, or Spark anywhere on the platform. There is also a real failure mode here: candidates over-invest in LeetCode because it is comfortable and gamified, then walk into a DE loop under-practiced on the things that actually decide it. Use it deliberately to clear the algorithm gate and to keep your raw coding sharp, then move the bulk of your hours to resources that target data engineering directly. Verdict: Necessary for the algorithm screen; thin for the data-engineering-specific rounds. 5. HackerRank HackerRank is where a surprising number of companies host their take-home and timed online assessments, so practicing in its environment carries a payoff most resources cannot offer: you get comfortable with the exact editor, the exact test-case runner, and the exact time-pressure UI you may actually be scored in. For an assessment you cannot retake, that familiarity is worth real points, because fighting an unfamiliar interface while the clock runs is a self-inflicted way to lose. Its SQL and problem-solving tracks are beginner-friendly, well-structured, and free to work through. The ceiling, though, is lower than you want for a senior DE loop. The problems lean academic and self-contained rather than job-realistic, the SQL rarely reaches the messy multi-table analytics that real interviews probe, and there is nothing on modeling, pipelines, or system design. The smart way to use HackerRank is as format rehearsal: run a few timed sets so the assessment environment feels routine, then build your actual depth somewhere that mirrors the work. Do not let a green checkmark on an easy problem set convince you that you are loop-ready. Verdict: Great for getting comfortable with the testing environment; limited depth. 6. SQLZoo A long-running, completely free interactive SQL tutorial that runs entirely in the browser with no signup, no setup, and no paywall. It walks you from SELECT basics through joins, grouping, subqueries, and window functions, with short hands-on exercises after each concept so you are writing real queries from the first lesson rather than just reading about them. For anyone whose SQL has gone rusty, or who learned it informally and has gaps they cannot quite name, it is the most painless way to rebuild muscle memory before stepping up to interview-grade problems. It is a teaching tool, not an interview platform, and you should treat it as exactly that. The problems stay introductory, the datasets are small and tidy, and there is nothing on data modeling, ETL, pipelines, or system design — the parts of the loop that actually separate data engineers from analysts. Its value is as a fast diagnostic and warm-up: work through the sections that feel shaky, confirm your fundamentals are solid, then graduate to harder, execution-based practice against realistic schemas. Linger here too long, and you will plateau well below where a real interview will push you. Verdict: A friendly free SQL primer; foundational rather than interview-level. 7. "Python for Data Analysis" (by Wes McKinney) Written by the creator of pandas, this is the reference for the kind of data-wrangling Python that shows up constantly in DE take-homes and pairing rounds: reshaping, grouping and aggregating, merging on imperfect keys, handling missing values, parsing dates, and cleaning the kind of messy tabular data that never looks like a tidy LeetCode input. Many data engineering interviews quietly assume this fluency, then hand you a notebook and a dirty CSV and watch how you move; if your Python is sharp on algorithms but clumsy on real data manipulation, this book is exactly the gap-closer. It is a library-and-technique book, not interview prep, and it will not touch SQL, data modeling, distributed compute, or system design. There are also no interview questions to grind, which is fine, because its job is to make the tools second nature so that during a timed exercise you are reasoning about the problem instead of fumbling for the right pandas idiom. Read the chapters on data loading, cleaning, and group operations, keep it nearby as a reference, then go apply the techniques in hands-on practice against problems that actually resemble the job. Verdict: The definitive practical Python reference for data work; not a drill book. 8. "Fundamentals of Data Engineering" (the book, by Joe Reis & Matt Housley) Another deliberate book pick, and the best single survey of the modern data engineering lifecycle: generation, ingestion, storage, transformation, and serving, plus the cross-cutting concerns like orchestration, data quality, and governance that interviewers increasingly probe. Where DDIA goes deep on systems internals, this book goes broad on how the pieces fit together into a working data platform, which is precisely the framing you want for the "walk me through how you'd build X" and "what would you consider before choosing this approach" portions of a loop. It is a framework-and-vocabulary book, not a practice book, and that is both its strength and its limit. It will give you the mental model and the shared language to discuss tradeoffs like a practitioner, which makes you sound, accurately, like someone who understands the field. But it contains no exercises, so reading it alone will not build the hands-on skill an interviewer also tests. Use it to organize everything you know into a coherent lifecycle, fill the conceptual gaps, then go write the queries and design the schemas somewhere that gives you real feedback. Verdict: The best lifecycle overview in print; conceptual, not hands-on. 9. Mode SQL Tutorial A free, well-regarded interactive SQL tutorial built by an analytics company, which shows in its framing: it teaches SQL the way analysts and engineers actually use it, oriented around answering real questions from data rather than solving abstract puzzles. It runs in the browser, takes you from the basics through intermediate analytics queries including aggregation and the early window-function territory, and the explanations are unusually clear about why a query is shaped the way it is. For someone shoring up SQL foundations before diving into harder problems, it is one of the cleanest no-cost on-ramps available. Like SQLZoo, it is a tutorial rather than an interview-prep platform, so it stops well short of the difficulty a real DE loop will throw at you, and it covers none of the modeling, pipeline, or system-design ground. It is best read as a companion to a hands-on platform: use Mode to internalize the analytical mindset and clean up your SQL fundamentals, then take that foundation into execution-based practice where the problems are harder, the schemas messier, and the feedback tells you exactly where your query went wrong. Verdict: A clean free SQL on-ramp; foundational rather than interview-level. 10. Pramp/Interviewing.io (mock interviews) Rounding out the list: peer and expert mock interviews. All the solo practice in the world cannot reproduce the specific pressure of explaining your reasoning out loud to a real human while a clock runs and someone is judging you, and that pressure is exactly where otherwise-prepared candidates fall apart. A handful of mock loops surface the weaknesses you cannot see in yourself: the long silences, the jumping to code before clarifying the question, the inability to narrate a tradeoff. Pramp pairs you with peers for free, while Interviewing.io connects you with experienced interviewers, often anonymously, for higher-fidelity feedback. The honest limitation is supply and specificity. Data-engineering-focused interviewers are scarcer than generalist software ones, so depending on availability, you may land in an algorithm or general system-design mock that only partially mirrors a true DE loop. That is still worth doing, because the communication skills, the structure, the clarifying questions, the calm narration, transfer directly regardless of the exact problem. Schedule one or two once your technical prep is underway, treat the feedback as data, and fix the delivery habits well before the interview that counts. Verdict: Best for rehearsing delivery and nerves; DE-specific matches can be hit-or-miss. How to Actually Use This List You do not need all ten. A focused plan beats a scattered one: Build the foundation. Skim CTCI for interview mechanics and start DDIA for concepts.Do the reps where it counts. Spend the bulk of your time on hands-on, DE-shaped practice that maps directly onto what you will be asked (see #1).Patch specific gaps. Use LeetCode for the algorithm screen, SQLZoo or the Mode tutorial to shore up SQL, and a mock interview or two to rehearse out loud. The candidates who get offers are not the ones who consumed the most content. They are the ones who practiced the actual job. Pick the resources that put you closest to it, start today, and write more queries than you read. Good luck with your loop.
Today, in modern backends, you probably have those distributed job queues for everything, including sending emails, processing payments, generating reports, and syncing data to third parties. As soon as you add retries to handle transient failures, however, you inherit a hard problem: how do you ensure that when the network, worker, or broker can fail at any point, your job runs exactly once? The short answer is: "exactly once delivery" is a great concept, but in practice it's mostly fiction given the nature of distributed systems. What you really can make is at-least-once delivery + idempotent processing, yielding exactly once effects. This article demonstrates how to accomplish this in Node.js with a tangible, functioning implementation. The Problem: Retries Cause Duplicates Take a worker that charges the customer and then marks the job completed TypeScript async function processJob(job) { await chargeCustomer(job.customerId, job.amount); await markJobComplete(job.id); } This seems fine until you consider that the work crashes after chargeCustomer succeeds but before markJobComplete executes. Because the queue does not receive an acknowledgement, it redelivers the job. The customer gets charged twice. This is not a rare edge case. Do any significant amount of throughput and workers fall over, containers reschedule, network calls default after the server has already worked its way through them. If you have side effects in your job, then you can always assume any job may be delivered more than once. The Solution: Idempotency Keys The main concept is to give every job created a unique, deterministic idempotency key and log the output of processing that key. The worker only checks if a key has been processed before doing any work. If so, it simply returns the result that was saved and does not redo the work. Here is the schema for how we can keep track of processed jobs. TypeScript CREATE TABLE processed_jobs ( idempotency_key VARCHAR(255) PRIMARY KEY, status VARCHAR(20) NOT NULL, -- 'in_progress' | 'completed' result JSONB, created_at TIMESTAMPTZ NOT NULL DEFAULT now(), completed_at TIMESTAMPTZ ); The job must encode its sensitive payload and key, not randomly generated at enqueue time. Good keys will be things like charge:order_12345, which will hopefully be stable across retries of the same logical operation. A Working Implementation The trick is to acquire the key atomically before doing anything useful. To claim the job, we execute a single atomic operation in PostgreSQL, which is an INSERT... ON CONFLICT DO NOTHING: TypeScript const { Pool } = require('pg'); const pool = new Pool(); async function processIdempotent(idempotencyKey, work) { const client = await pool.connect(); try { // Step 1: Try to claim the key atomically. const claim = await client.query( `INSERT INTO processed_jobs (idempotency_key, status) VALUES ($1, 'in_progress') ON CONFLICT (idempotency_key) DO NOTHING RETURNING idempotency_key`, [idempotencyKey] ); // Step 2: If we did NOT claim it, someone else already did. if (claim.rowCount === 0) { const existing = await client.query( `SELECT status, result FROM processed_jobs WHERE idempotency_key = $1`, [idempotencyKey] ); const row = existing.rows[0]; if (row.status === 'completed') { return row.result; // Return the cached result — no double work. } // Still in progress elsewhere — let the queue retry later. throw new Error('JOB_IN_PROGRESS'); } // Step 3: We own the key. Do the actual work. const result = await work(); // Step 4: Record the result. await client.query( `UPDATE processed_jobs SET status = 'completed', result = $2, completed_at = now() WHERE idempotency_key = $1`, [idempotencyKey, result] ); return result; } finally { client.release(); } } Now the worker becomes: TypeScript async function processJob(job) { return processIdempotent(`charge:${job.orderId}`, async () => { const charge = await chargeCustomer(job.customerId, job.amount); return { chargeId: charge.id }; }); } The key is already completed, and whenever this job is delivered the second (or more) time it will return the chargeId that was previously stored without charging again. Handling the Stuck "in_progress" Case The last failure mode that remains is where a worker picks a key, sets it to in_progress, and then dies without completing. Now this key is stuck, and any retry gives JOB_IN_PROGRESS forever. The solution is an expiration-lease for the lease. 1. Add locked_until column, make expired lock reclaimable: TypeScript const claim = await client.query( `INSERT INTO processed_jobs (idempotency_key, status, locked_until) VALUES ($1, 'in_progress', now() + interval '5 minutes') ON CONFLICT (idempotency_key) DO UPDATE SET locked_until = now() + interval '5 minutes', status = 'in_progress' WHERE processed_jobs.status = 'in_progress' AND processed_jobs.locked_until < now() RETURNING idempotency_key`, [idempotencyKey] ); It only requires a lock if the circuit is in progress and its lease has timed out, which means that some worker abandoned it earlier. The completed jobs will never be reclaimed, because the WHERE excludes them. Why not simply use a distributed lock One of the most common instincts here is to grab ourselves a Redis lock (if not using redis-lock, do SETNX with a TTL). Despite a lock being a solution for mutual exclusion, they do not solve idempotency by themselves. The job is already done, but because it uses a lock to prevent two workers from running at once. If you only use a lock, the job will be reprocessed when the lock expires and a redelivery is attempted. What you need is a permanent record of completion and that is what the processed jobs table provides. Locks and idempotency keys address two separate problems, yet durable systems typically require both. Takeaways Assume at-least-once delivery; make each job handler idempotent.Use the intent from job to derive idempotency keys; ensure they are stable across retries.Store results and claim keys atomically with INSERT... ON CONFLICT so duplicates return the cached result.Lease with an expiration because crashed workers should not block a key forever. Idempotency is certainly not useful, but it helps the queue to be the difference between something you can trust and a facility that will silently double charge your customers because of load. Treat it as a first-class citizen, because adding it via retrofitting after failing is way worse.
Most engineers learn these laws the hard way. When you try to rewrite something and it doesn’t deliver, or when a project is already late, adding engineers to the team will just make it fail faster. Sometimes, when you start using a metric to measure progress, the whole team will start trying to manipulate it. Then, six months later, someone mentions a 1975 law that addresses exactly what happened. I paid a price to learn this, too: I spent half my career learning these lessons the hard way, as many others probably did. The twenty laws listed below are the ones I refer to most often, although there are more (more on this later). Software development laws explain what is happening, what is about to happen, and what will not work no matter how hard you try. Some of these laws are sixty years old. They still apply to software development in 2026, and they will still apply in 2036 because they are not really about software. They are about people working together to build things under time pressure (basically, a lot of them are just laws of human nature). These laws are not rules that tell you what to do. They tell you what is already happening, but you still have to make the decisions. These laws just help you understand what is going on. Each of these laws made the list because I have experienced them myself. My book covers all fifty-six laws. If you only have time to remember twenty software development laws, these are the ones that I think are important. In particular, we will talk about the following laws: Gall’s Law: A complex system that works is always built from a simple system that worked first.KISS: Keep it simple. Anything beyond that is overhead.Conway’s Law: Organizations design systems that mirror their communication structure.Hyrum’s Law: With enough users, every observable behavior of your API becomes someone’s dependency, no matter what the contract says.CAP Theorem: A distributed system can guarantee only two of: consistency, availability, and partition tolerance.Zawinski’s Law: Every program expands until it can read mail. The ones that cannot are replaced by ones that can.Brooks’s Law: Adding people to a late software project makes it later.Ringelmann Effect: Individual output drops as team size goes up.Price’s Law: Half the work is done by the square root of the people.Dunning-Kruger Effect: The less you know about something, the more confident you tend to be.Hofstadter’s Law: It always takes longer than you expect, even when you account for Hofstadter’s Law.Parkinson’s Law: Work expands to fill the time available.Goodhart’s Law: When a measure becomes a target, it stops being a good measure.Gilb’s Law: Anything you need to quantify can be measured in some way that beats not measuring it.Knuth’s Optimization Principle: Premature optimization is the root of all evil.Amdahl’s Law: The speedup from parallelism is limited by the sequential part.Murphy’s Law: Anything that can go wrong will go wrong.Postel’s Law: Be conservative in what you send, liberal in what you accept.Sturgeon’s Law: 90% of everything is crap.Cunningham’s Law: The fastest way to get the right answer online is to post the wrong one. So, let’s dive in. How Systems Get Built 1. Gall’s Law A complex system that works is always built from a simple system that worked first. Systems do not work as well in real life as they do on paper because many problems do not surface until they hit the real world. These problems only appear when real users interact with systems, and by then, they either work or they do not. Every complex system that works got that way one step at a time. The systems that try to be perfect from the start usually fail. This is why most new versions of systems rewritten from scratch do not work out: teams keep all the features they had before, but lose the simple things that made the old systems good. Examples. Let’s take an example of Instagram. At the start, it was something else, but not a picture-sharing platform. The app was called Burbn, and it had: check-ins, gaming, photo sharing, all stuck together. Then, the founders cut everything except photo sharing, and the stripped-down core became the product. Google Wave went the other way. It launched with chat, email, a forum, and a document editor, all at once. Nobody could tell you what it was for, and it was dead in 15 months. 2. KISS (Keep It Simple, Stupid) Keep it simple. Anything beyond that is overhead. The KISS principle is a reminder that simplicity should be our key goal. If you can solve a problem with a 50-line script vs a complex 500-line solution, KISS favors the simpler solution because each line of code has the potential to cause an error. Why is simplicity so important? Software, in general, is complex to build and must be understood by humans. A simple design is much easier to maintain: new team members can get up to speed faster, bugs are easier to localize, and modifications cause fewer ripple effects. The KISS principle encourages developers to resist “clever” code that does too much at once, and to avoid architecting solutions that address future problems at the cost of current complexity. Example. Let’s say that we have a startup that needs a feature-flag system and decide to build a custom solution. They built it as a separate microservice with its own database, cache, admin UI, WebSocket notifications, and A/B testing support. It introduces a lot of complexity and takes a lot of time to build, which, if something goes wrong, can cause a lot of trouble. What they needed was a JSON config file. This would have taken an afternoon. 3. Conway’s Law Organizations design systems that mirror their communication structure. Your app architecture is already defined and essentially the same as your organization chart. For example, if you have four teams working on a project, you will probably end up with an app that has four parts. If the teams that work on the frontend, the backend, and the data do not communicate, your application will have three parts that do not work well together. If you rewrite your system without changing how your company is organized, you will still have the system, just written in a different language. The other way around works too. You can pick the architecture you want and then create teams that would naturally produce that kind of system. Amazon did this back in the 2000s. They broke their system down into smaller services managed by small teams, which changed how the system and the company worked together. This is called Inverse Conway’s Maneuver. Examples. Many modern AI organizations often split research from application engineering. Then, research optimizes benchmarks, while product ships apps against real users. The output is a model that scores well and a product that doesn’t work, because each side is optimizing for its own communication boundary. The pattern shows up at a small scale, too. A three-person team almost always ships a monolith because the cost of breaking it up is higher than the cost of keeping it together. 4. Hyrum’s Law With enough users, every observable behavior of your API becomes someone’s dependency, no matter what the contract says. The interface contract you wrote is not a proper contract. The real one is what your system actually does, including the parts you never expected to be important. For example, it could be timing, error message text, key order in JSON responses, and the exact bytes of a hash. Someone, somewhere, is depending on all of it. This is why backward compatibility costs so much in mature systems. This means that you actually don’t maintain the API you designed, but the accidental one. Examples. A good example is the SimCity game. I remember well that it had a use-after-free bug that worked fine on Windows 3.x because memory was never actually reclaimed. Then, Windows 95 reclaimed it, and SimCity crashed. Microsoft shipped Windows 95 with a special memory-allocator mode that was activated only when SimCity was running, so the bug would continue to work. Browsers do this at internet scale. Every quirk that web developers built into the platform effectively becomes part of it. The browser can’t change the quirk without breaking half the web. 5. CAP Theorem A distributed system can guarantee only two of the following: Consistency, Availability, and Partition tolerance. Networks fail. In a distributed system, that's not something you design around. It's something you accept. Once a partition happens, you have to pick: block writes to keep data consistent, or keep serving traffic and let replicas drift. Every distributed database makes this call. Most just don't tell you which one. They hide behind labels like "eventually consistent" or "highly available" and leave you to find out during an incident. Examples. MongoDB favors consistency, meaning that when a partition problem occurs, some MongoDB replicas will not accept any data until the entire system is working properly again. On the other hand, Cassandra will keep answering queries even when the replicas do not agree, and it will later fix the inconsistencies. Neither MongoDB nor Cassandra is wrong. They are just making choices about what your system can afford to lose. 6. Zawinski’s Law Every program expands until it can read mail. The ones that cannot are replaced by ones that can. Feature creep is not something that happens during the process. It is actually the process itself. When a tool is good at what it does, and people like it, they start using it all the time. The people in charge of the product want to keep the users engaged and stay on the platform. So the tool begins to take on tasks that are related to it. Over time, the tool becomes really slow and has a lot of unnecessary extra features. Then a new competitor comes along with a simpler version that does exactly the same thing. As the app's popularity grows, more and more unnecessary features are added. Examples. A famous example is Netscape, which started as a browser and ended as a suite with email, news, and a web editor. Firefox came as a fix and stripped it down, got popular, but then added plugins and a developer toolchain. We also remember Slack, which was launched to kill email and now has voice, video, bots, and an app directory. All of this is possible if the product doesn’t have the right north star metrics. How Teams Lose Speed 7. Brooks’s Law Adding people to a late software project makes it later. Software work is not easy to split among team members. When you bring someone new onto the project, it takes them a while to get up to speed, which means your experienced people have to stop what they are doing to help the new person learn. If your project is already behind schedule, adding more people won't make it go faster. It will just make things worse. Frederick P. Brooks said it well: you cannot have a baby in one month just because you have nine women pregnant. Software work is, like that, too. Software work does not get done faster just because you have people working on it. Example. Once, I was a team lead of eight people, and we were always behind schedule. My first thought was to hire two engineers to help us catch up. But in the meantime, while we were searching for new people, two people left us. It seemed that everything was now working better, communication was easier, and we managed to do more than before. So, obviously, the solution was to make the team smaller, not bigger. 8. Ringelmann Effect As teams grow, output per person falls. When many people pull on the rope, each person does not pull as hard. Some of this is because it is hard to work smoothly, and some of it is because people think someone else will do the part. Either way, this pattern is real. It is more extreme than most people think. Examples. A large GitHub study measured this directly. Developers on teams of 2-5 people averaged around 1,850 lines of code a month, while a team of 10 dropped to 1,200. At 50 or more, it was 450. Output per person fell 75%. This is why small teams ship faster than big ones, and why Amazon’s two-pizza rule holds true. It’s a defense against Ringelmann. This is especially true in today's AI-driven world, where productive teams have fewer members than before, as AI is driving up personal and team productivity. 9. Price’s Law Half the work is done by the square root of the people. In a group of 100 people, about 10 people actually do half of the work that matters. If you have a group of 16 people, it is likely that 4 people do most of the work. This is true for every creative field. The people in the group who do most of the work are really important, but the others are important too, because they do what needs to be done to support everyone else. They make sure everything runs properly (sometimes called glue work). So we need both groups, but the problem is that if the top people in your group leave, the group will lose a lot of its ability to get things done. Example. We all know that when Musk took over Twitter, it cut its staff by roughly 50%, and the site kept running. Price’s Law predicted that. What the law did not predict was what the layoffs removed: depth in trust and safety, SRE coverage, and incident response. The top performers kept the lights on. The organization lost the ability to handle the next hard problem, and Twitter quietly asked some laid-off people to come back. Why Plans Drift 10. Hofstadter’s Law It always takes longer than you expect, even when you account for Hofstadter’s Law. Let’s say you need to estimate how long something will take. You think four weeks is an estimate, but then you remember that your guesses are usually too optimistic, so you double it to eight weeks, just to be sure. But in the end, it takes sixteen weeks. Now you think, the next time you will be better, aren’t you? You think it will take sixteen weeks because that's what happened the last time. No, it now takes thirty-two weeks, because things you don’t know about surprise you. These are tasks such as unplanned integration issues or requirement changes. In practice, Hofstadter’s Law explains why techniques like padding estimates, awareness of Parkinson’s Law, and the use of historical data are essential, yet surprises still occur. Example. A good example of the Hofstadter law is the Berlin Brandenburg Airport project. The software integration process was taking much longer than expected, as it involved 75,000 sensors and 50,000 light fittings. The plan was to take 18 months to finish, but they later realized this was not possible and extended the timeline to 30 months. In the end, it took 7 years to complete, with a final cost of €7 billion. This was 2.5x higher than planned, and the airport opened 9 years late. 11. Dunning-Kruger Effect The less you know about something, the more confident you tend to be. Here is the uncomfortable part. The skill you need to do something is the same skill you need to judge how well you did the thing, and this is the problem. People who are not very good at something cannot see what they are doing wrong, so they think they are better at the thing than they really are. Yet, people who are good at it see all the things they are still getting wrong, so they think they are not as good at it as they really are. Examples. When asked when something will be done, new developers often give confident, precise estimates, while experienced developers give ranges (the famous “it depends” answer). The juniors aren’t wrong to be convinced. They simply don’t yet know what they don’t know (unknown-unknowns). People usually get really excited about new technology at first. This is because they have not used it a lot yet. We are seeing this happen with artificial intelligence now. The people who say AI can do anything are usually the ones who do not use it every day, like managers. 12. Parkinson’s Law Work expands to fill the time available. If you give a developer two weeks to do a task that can be done in two days, it will take two weeks to finish. This does not mean the developer is lazy or puts things off. People tend to fill up the time they have. Over the two weeks, the developer will likely spend time making plans, trying things, and adding extra tasks that do not need to be done (gold-plating). But if there was a deadline to have this done in a day, it would probably be done on that day. The thing about Parkinson’s Law is that it says if you give people a certain amount of time to do something, they will probably take all the time to do it. So, teams should set clear and realistic time limits (aka deadline-driven development). However, managers must use it judiciously, combining Parkinson’s insight with realistic scheduling. If you compress timelines too much, you risk running into Hofstadter’s Law, which reminds us that work often still takes longer than expected, even with buffers. Examples. A developer given two months for a one-week task will spend a month prototyping alternatives, another week on architecture debates, and the last three weeks polishing details nobody asked for. If we give the same task, but this time with a clear one-week deadline, it will be shipped in one week. How Metrics Distort Work 13. Goodhart’s Law When a measure becomes a target, it stops being a good measure. We can use many different ways to measure our work, e.g., number of bugs closed, number of incidents, test coverage, or team velocity. When we start measuring people's performance based on these things, they will focus on making those numbers look good instead of actually doing good work. The numbers will go up, but the work will not get any better. This is because when we give people incentives, they will do what gets them the reward, not what we really want. When we measure the wrong thing, people will do the wrong thing to get ahead. Examples. I watched a team get rewarded for lines of code written at the start of 2000, and the number of PRs created some years later. Developers started copy-pasting instead of extracting shared logic. Some created PRs for almost every commit they made. The modern version is AI tokens consumed per engineer (called tokenmaxxing). More tokens are being treated as a sign of productivity. 14. Gilb’s Law Anything you need to quantify can be measured in some way that beats not measuring it at all. Gilb's Law is like the side of the coin to Goodhart’s Law. You can say, when looking at Goodhart’s Law, that having metrics is bad, but that is actually not true. Not having any metrics is even worse than that. If something is important to you, you should try to find a way to measure it, because we cannot improve what we don’t measure (as Peter Drucker famously said). Example. Developer productivity is usually a hard thing to measure, and it always has been. We had many bad metrics, from lines of code to token consumption. But deployment frequency and change lead time give you a signal (as in the DORA metrics for DevOps) as a proxy. What Breaks Under Load 15. Knuth’s Optimization Principle Premature optimization is the root of all evil. Most performance work happens too early and in the wrong place. Teams optimize code paths that never become hot, introduce complexity they never need, and burn time solving a scale problem they may never earn. So the best way is to write the code that works, then check its performance. If there is a problem, a tool will show you where it is. If not, just move on. Examples. I worked at a startup once, where we spent a lot of time setting up Kubernetes. The thing was that we did it to handle millions of users, and we didn’t even have 10 users yet. We were making our infrastructure ready for a load that didn’t exist. Our product features were not even finished. One of my colleagues said that we should make sure 100 people even want our product before we worry about handling millions of users. He was right. We still launched late. 16. Amdahl’s Law The speedup from parallelism is limited by the sequential part. If 10% of your work has to be done in a sequential way, the work will only go 10x faster, no matter how many computers you use. If 50% of the work has to be done one thing at a time, the work will only go twice as fast. The same thing happens with people. If one group of people has to say yes to every decision, about how something is built, that limits how fast your team can work, no matter how many engineers you have. If you add engineers, but they all have to wait for the same group of people to say yes, the line of people waiting just gets longer. Your team of engineers will still be slow because the group of people making decisions is a bottleneck. The work of your team of engineers will only go as fast as the group of people making decisions. Examples. Scaling web traffic by adding more app servers helps until every request hits one shared database or authentication service. Then adding more horizontal scaling doesn’t help. The conversation about AI productivity is hitting the roof now. AI makes coding faster, but you still have to think, check, fix errors, and work together on those steps that can’t be done simultaneously. This sets the limit on how much you can gain in the end. That’s why some engineers see their work speed up by 10 times, and others see a 1.2 times increase. 17. Murphy’s Law Anything that can go wrong will go wrong. In software, Murphy’s Law is often mentioned to explain bugs and production incidents: whatever can go wrong in code (a null pointer, a race condition, a network outage) will eventually manifest, especially in large user bases or at the worst possible time (Friday evening). In practice, this law encourages developers to write more defensive code. This means checking for nulls, handling exceptions, validating inputs, and failing gracefully when errors occur. It also reminds DevOps teams to anticipate failures by implementing monitoring, enabling rollbacks, and maintaining contingency plans. Example. On July 19 2024, CrowdStrike made a change to the Falcon Sensor settings. This change caused a memory issue on Windows machines. It made 8.5 million Windows machines stop working and show a screen. To fix this problem, someone had to log in to each machine and apply the fix, because those machines could not start up. This could be done remotely. And this happened on a Friday morning when no IT staff members were working. It caused problems for airlines, hospitals, and banks. Everything that could go wrong did go wrong on the day, just like Murphy’s Law says. 18. Postel’s Law Be conservative in what you send, liberal in what you accept. This law says that if your server sends HTTP responses, it should format headers exactly per spec. But if your server receives an HTTP request with an uncommon header order or an unusual format, you should still process it rather than drop the connection, as long as you can interpret it safely. Browsers do this at a scale. Most of the HTML on the web is not written correctly, but modern browsers still render it. If they were strict, half the internet would not be found. But there is one thing to consider. Being too liberal has a cost: if everyone accepts anything, problems will never be corrected. There will be just more mess. In security-sensitive code, tolerating input can make it easier for attackers to find. So, the basic idea still holds. You need to use judgment, as being lenient is not the same as being permissive. Example. In APIs, say your service expects a timestamp. If it receives a timestamp without a time zone, instead of rejecting, maybe you assume UTC or try to parse it anyway, being liberal in acceptance. But when your service returns data, you always include the time zone to ensure the output is conservative and precise. How to Judge Better 19. Sturgeon’s Law 90% of everything is crap. Most things we make will go unused, and most of the code we write is not good. Most projects we start do not deliver the value that we thought they would. This is not a bad thing per se. This is how things are when we are trying to create something new. If we pretend everything is great, we will treat every project the same, which will make things too complicated. The projects that really matter are the ones, like 10% of them. Finding these projects and getting rid of all the others is what really takes skill. Example. WordPress has roughly 57,000 plugins in its directory. Over 34,000 haven’t been updated in the past 2 years, and nearly 19% have zero active installs. A small number of well-maintained plugins powers 40%+ of the public web. That distribution is Sturgeon’s Law in one screenshot. 20. Cunningham’s Law The fastest way to get the right answer online is to post the wrong one. When you ask a question on some online forum, you usually get no response. If you post something that is clearly incorrect, people will jump in to correct you. They might just walk by if they see a question, and then cannot help themselves when they see something that is wrong. You can actually use this to your advantage. If you are having trouble with something, do not ask how you should do it. Instead, propose a solution you know is not very good, or share a draft, and then see what happens. The right answer might come to you without you even asking for it. Note that this trick only works when the people around you know what they are talking about. If you are in a group where everyone’s just as confused as you are, then a wrong answer can actually cause more harm than good. In that case, the wrong answer can just become information that people start to believe. Example. The whole bet of wikis, and later Wikipedia, runs on this insight. People correct errors faster than they write articles from scratch. The bet paid off on a civilization-scale. Conclusion In this article, I shared some of the most impactful laws I saw in my career. You do not have to memorize all of them. The top five or six laws will help you solve most of your issues. The rest are there for when a new problem arises. What is more important is knowing when a law applies and when it does not. These twenty laws often conflict with each other. Knuth says do not optimize early. Amdahl says find and fix the part of your project that is slowing everything down. Both are correct at times. The key is to know which one to use now. Also, this list is my list. Your list will be different. The laws that have caused you problems will be more important to you than the ones that have not. Over time, you will add your laws. Write them down when you notice them. One line per project, incident, or rewrite. Which law helped you? Which law gave you advice? What changed? Your personal list will be more helpful to you than any list I can give you. Frameworks, platforms, and deployment models have changed since Brooks wrote his book in 1975. These laws have not changed. They describe the one thing that has not changed: humans building things together under constraints they do not yet fully understand. That is why they are worth learning before the project, not after it causes problems.
She had everything on the list. Eight years of experience. Strong systems design. Distributed architecture under her belt. The panel interview went well — one of the hiring managers later described it as the best technical conversation they'd had with a candidate all quarter. The team passed on her. Two weeks later, during a casual conversation with that hiring manager, the reason came out. It wasn't her architectural skills or her communication. It was a question someone had slipped in near the end: "Walk us through how you'd set up an AI-assisted code review pipeline for a team that ships twelve microservices." She described doing it manually. The other finalist described standing up an orchestration layer with context-aware models, configuring fallback thresholds, and building observable feedback loops that trained the team's prompt library over time. Same job title. Completely different mental model of what the job now involves. That story isn't unique. It captures something that's been happening gradually over the past eighteen months and then very suddenly in the last six: the senior developer role has quietly split into two jobs. One of them is the job we all trained for. The other is the job that a meaningful portion of your working week now actually requires. And the gap between developers who've accepted that and developers who haven't is becoming very hard to explain away in performance conversations. The Split That Happened Without a Memo Let's be specific about what the "AI Systems Architect" half of the role actually means, because people either over-mystify it or undersell it. It doesn't mean you become a data scientist. It doesn't mean you're fine-tuning models or writing PyTorch. Those are real jobs — they're just different jobs. What it means is something more operational and less glamorous: you are now responsible for designing, maintaining, and improving the systems of AI assistance that your team works inside of, not just the code that the team produces. That sounds abstract until you break it into daily decisions. Which tasks should be fully AI-generated versus AI-assisted versus AI-reviewed only? Where are your model's blind spots for your specific codebase, and how do you account for them in code review? When a junior developer on your team gets a plausible-but-wrong architectural suggestion from an AI assistant, what's the escalation path? How do you measure the quality of your team's prompting over time? These aren't rhetorical questions — they're operational ones that live teams are answering right now, often badly, because no one assigned anyone to own them. Senior developers are getting assigned to own them. Not officially. Not with updated job descriptions. Just through the ordinary mechanism of "this problem needs solving, and you're the most experienced technical person in the room." What "AI Systems Architect" Actually Means Day to Day The phrase sounds bigger than the practice. What it actually breaks down to is four interconnected responsibilities that are now landing on senior developers, whether they want them or not. First: workflow design. Someone has to decide which parts of the development cycle use AI assistance, at what level of autonomy, and with what human checkpoints. At most companies, this currently happens by accident — everyone develops their own habits, and nobody compares notes. The developers who are stepping into the architect half of the role are the ones making that deliberate, rather than emergent. Second: model selection and configuration. Not fine-tuning, but product-level decisions: which models for which tasks, what context window strategy, how to handle codebases that exceed context limits, what fallback behavior looks like. These are practical engineering decisions that live in the space between "developer tool choice" and "infrastructure decision." They belong to senior engineers. Third: quality governance. AI-generated code introduces a new failure mode: plausible-looking outputs that are subtly wrong. The patterns of wrongness are specific and learnable. Senior developers who have mapped the failure modes of their AI tooling — the kinds of edge cases it consistently misses, the naming convention assumptions it gets backward, the security patterns it handles confidently and incorrectly — are providing a form of institutional knowledge that is genuinely hard to replace. Fourth: team prompting culture. This is the one nobody talks about at conferences yet, but engineering managers across the industry have been mentioning it consistently over the past six months: the quality variance in how different team members prompt their AI tools is enormous, and it compounds. Senior developers who build and maintain shared prompt libraries, who do prompt review the way they do code review, who can diagnose why a colleague got a bad output — those developers are operating as a force multiplier for the entire team, not just themselves. The Job Description Before and After: A Concrete Comparison This is worth making explicit. Analysis of actual senior engineer job postings — anonymized, from companies between 80 and 1,200 employees — shows a clear shift when comparing what the role requirements looked like in early 2023 versus what's being written now. The change is real and measurable. The pattern across all of it: the what of the role hasn't changed so much as the how and the governance around it. Senior developers are still responsible for the same categories of work. They're now also responsible for the design of the AI-assisted systems that help a team do that work, and for the failure modes those systems introduce. The New Core Competency Stack Here's what the competency model looks like in practice when you lay it out. The traditional side should feel familiar. The AI architecture side probably contains a few items you haven't formally owned yet — but if you've been doing this job for more than two years and paying attention, you've been building these skills without realizing it. The Salary Premium Is Already Real Compensation data lags reality by about eighteen months, so take specific numbers here with appropriate skepticism. What industry reporting suggests is that a clear pattern is emerging: developers who can demonstrably operate in both halves of the new role — not just use AI tools personally, but architect AI-assisted workflows for a team — are commanding a premium that's running somewhere between 18% and 31% above their single-track counterparts at the same years-of-experience mark. That range is wide. The premium is highest in companies that have recently invested in AI transformation initiatives and learned, the hard way, that "everyone uses Copilot" is not the same as "we have a coherent AI engineering strategy." Those companies are specifically recruiting for systems architect skills because they've already paid for the gap. How to Build the Second Half of the Job Nobody teaches this in a course yet. There are some good books and a growing number of blog posts, but the skills are mostly developed through deliberate practice and iteration. Based on teams that have successfully made this transition, here's what works. The starting point is mapping your team's current AI-assisted work honestly. Not aspirationally — honestly. Which tasks are you and your team currently doing with AI assistance? Where does the output go without sufficient review? What are the categories of error you've caught, and what categories might you be missing? This audit, done once and updated quarterly, is the foundation of a governance practice. From there, the most leveraged thing most senior developers can do is build a shared prompt library for their most common task types. Not a personal one — a shared one, with a versioning and review practice attached. The discipline of reviewing a colleague's prompt and explaining why it produced a wrong output is one of the fastest ways to build the mental model you need for the governance half of the role.
AWS has been building agentic infrastructure for some time now — Bedrock, AgentCore, Strands — mostly aimed at engineers who want to build their own agent systems from scratch. Amazon Quick is a different layer of the same bet: a ready-to-use agentic workspace that targets teams directly, without requiring custom orchestration code. This article walks through what Quick is, how its components fit together technically, how the MCP integration model works with real code, and where it sits relative to the rest of AWS's agent stack. What Amazon Quick Is Amazon Quick is an AI assistant for work that connects to your existing tools — Slack, Microsoft Teams, Outlook, CRMs, databases, and local files — and gives a unified layer for querying, automating, and acting across them. It launched in preview at AWS's "What's Next with AWS" event on April 28, 2026. The product is aimed at teams, not just individual users. One person can build a custom agent scoped to a specific dataset or workflow, and the whole team benefits from it. Responses from Quick agents are grounded in your actual business data, not the underlying model's training distribution. Under the hood, Quick is built on Amazon Bedrock AgentCore and uses the Model Context Protocol (MCP) as its standard for connecting to external tools. It runs on AWS IAM and VPC, which means it inherits the same security and compliance posture as the rest of your AWS workloads. Components Quick bundles five distinct capabilities. It helps to understand each one separately before thinking about how they compose. ComponentWhat it doesSpacesCollaborative workspaces where teams pool files, dashboards, and data sources. Agents in a Space are grounded in that Space's data.AgentsCustom, domain-scoped agents built on your team's specific data. One person builds, everyone uses.ResearchMulti-source synthesis across internal data, the public web, and third-party datasets. Produces structured reports.Visualize (Quick Sight)Integrated BI layer. Conversational access to dashboards, charts, and forecasting — no separate BI tool required.Automate (Quick Flows)Workflow automation from simple daily tasks to complex multi-step processes with cross-app action execution. Each component is available through the web app, mobile, and a native desktop app (currently in preview for macOS and Windows) that can read local files and calendar context without requiring browser access. Where Quick Sits in the AWS Agent Stack AWS is building in two directions at once. AgentCore is the infrastructure layer for engineers who want to compose their own agent systems — runtime, memory, gateway, observability — with any model and any framework. Quick is the product layer on top: opinionated, team-facing, and deployable without writing orchestration code. The practical implication: if you're an engineer building internal tools or automation pipelines, you'll likely interact with both layers. AgentCore for the infrastructure wiring; Quick as a surface where non-technical teammates interact with the agents you build. The Integration Architecture The core question for any engineer evaluating Quick is: how does it actually connect to external systems, and what does the request path look like? Quick uses MCP (Model Context Protocol) as its primary integration standard. This is significant because MCP is an open protocol — it means Quick agents are not locked into AWS-specific connectors, and any MCP-compatible server can be registered as a tool source. High-Level Request Flow The sequence below shows the full lifecycle of a single agent-triggered tool call — from the moment Quick receives a prompt through to the response returning from a downstream API. Quick acts as the MCP client. Your MCP server exposes tools via listTools and callTool. Quick discovers them at registration time and makes them available to any agent or automation in the workspace. Authentication flows through OAuth 2.0, with support for Dynamic Client Registration (DCR) so Quick can register itself automatically without manual credential setup. Building an MCP Server for Quick Here is a minimal Python MCP server using the mcp SDK that exposes two tools Quick can invoke — get_ticket and list_open_tickets. This pattern works whether you host the server yourself or run it on AgentCore Runtime. Install Dependencies Python pip install mcp[server] httpx uvicorn Server Implementation Python # server.py from mcp.server import Server from mcp.server.sse import SseServerTransport from mcp.types import Tool, TextContent import httpx import json from starlette.applications import Starlette from starlette.routing import Route app = Server("jira-quick-integration") JIRA_BASE_URL = "https://yourorg.atlassian.net" JIRA_TOKEN = "Bearer
VP of Engineering,
Factorial
CTO,
Multiplayer
Owner,
AlexVakulov
Chief Technical Architect / Fractional CTO,
NIQ