Synthesis and timing closure¶
Think in hardware¶
Before writing RTL, identify:
- Registers.
- Combinational operations between registers.
- Data widths.
- Clock frequency.
- Latency and throughput.
- External interfaces.
A long expression may infer a long combinational path. The source being one line does not make the hardware fast.
Resource inference¶
Vivado can infer:
- Flip-flops and LUTs.
- Carry chains for arithmetic.
- Distributed RAM and block RAM.
- DSP slices for multiply/accumulate.
- Shift-register LUTs.
- Clock enable and reset pins where mappings allow.
Inference depends on coding style, widths, reset behavior, and target device. Confirm the result in synthesis reports.
Pipelining¶
Pipelining inserts registers to reduce combinational delay per cycle:
un-pipelined: register → large calculation → register
pipelined: register → part A → register → part B → register
Tradeoffs:
- Higher maximum clock frequency.
- Increased latency.
- More registers and control alignment.
- Throughput can remain one result per cycle after filling.
Update the testbench scoreboard latency whenever the pipeline changes.
Fanout¶
A reset, enable, or control net driving many destinations may have large routing delay. Remedies can include:
- Register replication.
- Hierarchical/local enable generation.
- Pipelining.
- Letting Vivado physical optimization replicate drivers.
- Reducing unnecessary global reset use.
Do not manually duplicate logic without checking the reports and functional effect.
Reading a timing path¶
For the worst path, determine:
- Launch clock and source register.
- Destination register and capture clock.
- Logic levels.
- Net/routing delay.
- Clock skew and uncertainty.
- Required time, arrival time, and slack.
Then decide whether the problem is:
- Too much logic: pipeline or restructure.
- Too much routing: reduce fanout, adjust hierarchy, or improve placement.
- Incorrect clock relationship: fix constraints/CDC architecture.
- Unrealistic requirement: revisit specification knowingly.
Setup versus hold¶
Setup failures are often improved by reducing maximum data-path delay or allowing more legitimate cycles.
Hold failures concern minimum delay. Do not “fix” them by changing RTL blindly; implementation tools usually insert route delay, and incorrect clock/exception constraints may be the actual problem.
Latency, throughput, frequency¶
Keep these distinct:
- Latency: time/cycles from input acceptance to corresponding output.
- Throughput: transactions accepted/completed per unit time.
- Clock frequency: clock edges per second.
A five-stage pipeline may have five-cycle latency yet accept one input every cycle.
Utilization¶
Inspect:
- LUTs as logic and memory.
- Flip-flops.
- Block RAM.
- DSPs.
- I/O pins.
- Clocking resources.
Unexpected resource counts are diagnostic. Examples:
- Thousands of flip-flops from an accidentally unrolled structure.
- LUT RAM instead of block RAM due to reset/read style.
- DSPs not inferred because operands/types do not match an inference pattern.
Synthesis warnings¶
Classify each warning:
- Expected and documented.
- Harmless for the target but understood.
- Indicates a real bug.
- Unknown—investigate before proceeding.
Particularly important:
- Inferred latches.
- Multiple drivers.
- Width truncation.
- Unconnected or undriven ports.
- Removed registers/logic.
- Combinational loops.
- Clock used as data or data used as clock.
Timing-closure workflow¶
- Verify all clocks and I/O constraints.
- Check for unconstrained paths.
- Classify cross-clock paths.
- Identify worst failing path groups.
- Fix architectural problems first.
- Re-run implementation.
- Compare reports; do not rely on intuition alone.
- Preserve the testbench and functional behavior during optimization.
Definition of timing success¶
Timing closure means all intended paths satisfy their constraints in the implemented design. It is not:
- Simulation at 10 ns.
- Synthesis completing.
- Bitstream generation completing.
- One favorable implementation run while other required path groups fail.