Sunday, May 3, 2020

Brushed DC motor driving and sensorless speed feedback

Introduction


A brushed DC motor seems like a very simple thing. You connect a voltage on the terminals, and it spins. If you reverse the polarity, it spins in the opposite direction. However, I've learned the whole picture is not quite so simple. There are all sorts of weird and wonderful things occurring when driving the motor with pulse width modulation.

Some time ago I had a need to control DC motors in speed feedback, but without the ability to add any sensors to the rotor to measure the speed directly. Turns out I'm not the first to have this situation in my hands, and not the first to try and figure out how to do it. Precision Microdrives has a good article on the subject. However, for my application I needed bidirectional motor control with an H-bridge, which their article does not cover.

I've tried to capture here what I've learned. Hopefully useful for others, but at least as a reference for myself.

Brushed motors 101


A typical cheap brushed DC motor has a permanent magnetic field in the stator and a rotating armature with three rotor coils.
Construction of a typical toy DC motor. Image by Haade from Wikimedia Commons

A current through the motor terminals will pass through to the armature via brushes, which brush against the commutators. The commutators are connected to the three coils in such a way as to produce a torque in one direction, regardless of the orientation of the rotor. The basic idea behind this is very well illustrated in this YouTube video.

It is often assumed that the commutation event, when a brush changes from one commutator to another is instantaneous. In reality the coils have inductance, and thus instantaneous changes in current are not possible. I've not found a proper explanation of the specifics happening during this event. It may be that the residual magnetic energy simply gets dissipated through arcing (as seen in this video, although that motor has a much greater number of poles). In any case, the fact that there is significant inductance in the armature makes DC motors much more interesting.

The ideal DC motor equivalent electrical circuit is simply a voltage source (representing the so called back-EMF of the motor) with a voltage that is directly proportional to the rotation speed of the motor (the KV factor). That is, the voltage across the motor, when rotating at speed omega is


Spoiler alert: in the end it is this voltage, which allows us the measure the rotation speed without sensors.

The torque M produced by the ideal motor, on the other hand, is directly proportional to the current I passing through the armature (The KT factor). That is,


Now, the electrical power going in to the motor is Vemf*I, while the mechanical power is


In the ideal case the mechanical power matches the electrical power, and thus (when using SI units, omega in rad/s)
See also Wikipedia.

However, connecting an ideal motor to a voltage source would show infinite current, which is clearly not correct. To improve the model, an obvious thing to add is the wire resistance of the armature, which is seen in series with the ideal motor. This already improves the model, and is enough to handle the static situation. Unfortunately it still fails to catch the dynamic effects, when the motor voltage is switched quickly high and low, as happens in an H-bridge with pulse width modulation. To catch this dynamic behavior, the winding inductance also needs to be accounted for.

Ideal motor model with armature resistance and inductance added.

Motor driving modes


As mentioned before, the plan is to drive the motor with an H-bridge driver. This allows driving the motor in both directions, and the speed can be varied by pulsing the motor. However, there are a few ways a motor can be driven in an H-bridge.

Below are the typical four configurations used by common motor driver ICs. Illustrated current is the steady state current, after inductive dynamic effects have died out.
Driving the motor in the forward direction. Motor positive is connected to VCC, motor negative is connected to GND.
Driving the motor in the backward direction. Motor positive is connected to GND, motor negative is connected to VCC.
Braking the motor. Both motor terminals are connected to GND. Current flow direction depends on the motor rotation direction.
Motor freewheeling. Both motor terminals are floating. There is no current flow and motor terminal voltage is proportional to rotation speed.
When a motor speed is controlled using PWM in an H-bridge, the motor is either pulsed between the drive and brake configurations or between the drive and freewheel configurations. My intuition intially was that pulsing the motor with the brake configuration would be wasteful, as it seems obvious that braking would dissipate mechanical energy as heat, which reduces the efficiency.

Let's take a closer look at the two different pulsing strategies. First, let's lay down some hypothetical model parameters of the motor. Take VCC as 5V, and say the motor is rotating at a speed, which produces 2.5V as back-EMF, and say the motor takes an average current of 0.1A to maintain this speed. Take the armature resistance as 0.1 ohms and the armature inductance as 10 millihenry. Assume everything we consider happens so quickly, that the motor doesn't have time to change its speed.

That is


Forward drive

Motor driven in the forward direction. Current flow indicated with red arrow.
Initially we're in the forward drive configuration. Motor positive terminal is connected to 5V and negative terminal is connected to GND. This drives a current I (approximately Iavg) through the motor. Writing the voltage drop in the circuit, we obtain the equation (omitting the H-bridge switch resistances)

We see that the current is in fact not constant, but changes in time due to the inductance. The change is

With our hypothetical values, this evaluates as I' = 249 A/s. That is, the current is increasing in time and the armature magnetic field is growing.

Let's assume that the motor pulsing scheme is such, that the current doesn't grow or decrease much. Thus the current I remains approximately Iavg at all times. If the switching frequency were e.g. 10 kHz, then the current would change a maximum of only 0.0249A during a cycle in our example.

Multiplying the equation for the voltage drop with the current (and arranging a bit) gives the power balance as


The terms can now be identified as
  • Vcc*I is the input electical power
  • Larm*I'*I is the power getting stored in the growing magnetic field of the armature
  • Rarm*I² is the power dissipated in the armature resistance
  • Vemf*I is the power getting converted to mechanical power
Turns out that the only term where power is lost is the armature resistance.

Braking (also known as slow decay mode)

Motor switched to brake configuration, while current is flowing. Armature inductance forces current to flow. Current flow indicated with red arrow.
Next, we consider what happens when the H-bridge switches from the drive configuration to the brake configuration. Here we have both terminals of the motor connected to GND. The voltage drops around the circuit give the following equation

Again, we see current is not strictly constant in time. However, it must remain continuous over the switching event. Solving for the current derivative now gives

Which with our hypothetical values evaluates as I' = -251 A/s. That is, the current is decreasing with time and the armature magnetic field is collapsing.

Again, multiplying the equation for the voltage drops with the current (and re-arranging) gives the power balance as

The terms can again be identified as
  • -Larm*I'*I is the power extracted from the collapsing magnetic field of the armature
  • Rarm*I² is the power dissipated in the armature resistance
  • Vemf*I is the power getting converted to mechanical power
Turns out again that the only term where power is lost is the armature resistance. So in a short time scale the braking isn't in fact dissipating any mechanical power.

If we stay in the brake configuration for a long period of time, eventually the current will decrease to zero, and go negative. At this moment the torque produced by the motor to the mechanical load reverses, and the motor starts to take energy from the mechanical side back to the electrical side, storing it to the armature magnetic field, which now is growing in the opposite polarity as the inital drive was.

Freewheel (also known as fast decay mode)


Knowing now, that the current through the motor cannot change instantaneously due to the inductance, it is not immediately clear what will happen if the H-bridge were to switch to the freewheel configuration from the drive configuration. In fact, the H-bridge needs to be carefully designed to be able to handle such a situation.
Motor switched to freewheel configuration, while current is flowing. The armature inductance forces current through the parasitic diodes of the MOSFETs. Current flow indicated with red arrow.

When the H-bridge is built using MOSFET transistors (as they typically are), there are (parasitic) body diodes inside the MOSFETs, which start to conduct, allowing the current to continue flowing as shown above. Often there will be additional Schottky diodes parallel to the MOSFETs to not just rely on the body diodes.
Motor switched to freewheel configuration, while current is flowing. The armature inductance forces current to flow. Current flow indicated with red arrow.
The voltage drops around the circuit gives us

Solving for the current derivative now gives

With our hypothetical values (and Vdiode = 0.7V), this evaluates as I' = -891 A/s. The current is decreasing, and thus the magnetic field is again collapsing. Since the current is decreasing much faster here than in the braking configuration, this motivates the alternate terms "slow decay" and "fast decay" for these two configurations.

We again get the the power balance through multiplying the equation for voltage drop with the current

The terms are now identified as
  • -Larm*I'*I is the power extracted from the collapsing magnetic field of the armature
  • Vdiode*I is the power dissipated in each diode
  • Rarm*I² is the power dissipated in the armature resistance
  • Vemf*I is the power converted to mechanical power
  • Vcc*I is power that is absorbed back to the power supply
In addition to the armature resistance, the diode drops also contribute to the power loss in this configuration. Further, depending on the type of power supply, the current fed to the supply rail may also constitute a power loss. Thus in a short time scale, the freewheel configuration will be less efficient. Additionally, since the decay rate of the current is also much greater, a larger duty cycle in the drive configuration is needed to keep the average current at the desired value.

If we were to remain in the freewheel configuration for a long time, the current would decrease quite quickly to zero. However, since the current is passing through diodes, the current can not change direction and will simply stop when the current reaches zero. Since the current is then zero and also unchanging, the rotation-speed-dependent back-EMF of the motor can be seen and measured across the terminals. This is the key for the sensorless speed measurement.

Driving scheme for sensorless speed feedback


The mechanical dynamics of the systems are much slower than the electrical dynamics of the system. That is, the motor rotational speed remains constant in the time scale of say 100 Hz, while the time scale to consider the current a constant is around 10 kHz. It is enough to measure the motor speed at 100 Hz, while the switching occurs at 10 kHz.

Since the drive-brake pulsing is more efficient than drive-freewheel, it makes sense to use that most of the time, while occasionally going to the freewheel configuration and letting the current fully decay in order to measure the speed.

Two important things to note regarding the motor current decaying.
  1. The motor torque remains directly proportional to the current. Having the current decay to zero and back up again causes torque oscillations, which may cause mechanicals problems if the frequency is close to a mechanical resonance.
  2. The current decay rate is a function of the back-EMF of the motor. When the motor is turning quickly, the current decays up to twice as fast as to when the motor is turning very slowly. Further, the decay time is also proportional to the motor current at the start of the decay. If a fixed delay is used to wait for the current decay, it needs to be determined by stalling the motor at the maximum expected current.
The measurement of the back-EMF is very noisy. A significant source of noise being the commutation events, which will occasionally occur during measurement. It is a good idea to take a few measurements during each current decay period and pick the median. Also, the conversion from back-EMF to rotational speed is not exact and varies from coil to coil and also on the orientation of the coils to the stator field.

All that said, if torque oscillations and motor drive efficiency are not relevant issues, perhaps the simplest solution is to run the motor with freewheel pulsing using a pulse frequency low enough, which results in the current decaying to zero in every cycle. This allows also measuring the back-EMF in every cycle, which then permits reducing noise through a simple low-pass filter.

Wednesday, December 4, 2019

GameBoy to VGA

VGA mods for the GameBoy DMG-01 aren't anything particularly new. There are several such projects laying around the internet. BennVenn even used to sell a mod kit.

My main motivation, however, is not to just have a VGA out on my GameBoy, but to teach myself Verilog through a non-trivial project. For development, I started with a cheap Cyclone 1 (EP1C3) based FPGA board I had ordered previously.

Hardware

The board, labeled "FPGA NIOS KIT Board TodayStart" has a two big bugs, and I can't really recommend this board for this use. The 50 MHz oscillator is connected to CLK3, which is not connected to the PLL on the EP1C3. Also, the PLL supply pin is connected to 3.3V through a 100 ohm resistor, which causes the PLL to not work properly even if you manage to connect a clock to it.

I had to first fix the issues by soldering a wire from the oscillator to CLK0 and by replacing the 100 ohm resistor with a 1 ohm one.

Some further discussion of the board (in German) at mikrokontroller.net forums.

The dev board I started the project with.

The display resolution on the GameBoy is 160x144 pixels, with 4 shades of gray per pixel. The pixel aspect ratio is 1:1 and the display is updated at approximately 60 Hz. Although the refresh rate matches VGA, no VGA resolution has this few pixels. Some type of upscaling is required, and thus also some form of image buffering. The EP1C3 has 59904 bits of block memory, which is just enough for buffering one GB frame (160*144*2 bits = 46080 bits).

Scaling each dimension by 4 results in a scaled image of size 640x576. A common resolution close to this is the Super VGA at 800x600. Using an 800x600 output and 4x scaling the GB frame, we are left with 80 pixels padding on the left and right side of the image, but only 12 pixels on the top and bottom. Since the image aspect ratios do not match, we obviously can't get rid of the horizontal padding completely. With this choice, however, the vertical padding is almost as little as possible.

The hardware to connect the FPGA to the monitor couldn't be much simpler. Sync signals use TTL logic levels and thus the 3.3V outputs from the FPGA are directly compatible. As for the video signal itself, I built a two bit DAC out of two resistors for each color channel. Taking the resistor values of 390 and 820 ohms together with the 75 ohm input impedance of the RGB inputs, gives the possible output voltages of 0V, 0.24V, 0.49V and 0.73V. This matches quite well the nominal signal levels of 0-0.7V specified for the RGB inputs.

The LCD drive signals from the GameBoy use 5V CMOS logic, which is incompatible with the FPGA inputs. For now I am using simple resistor divider level converters to get the 5V down to 3.3V. The resistor values I use are 2.2k and 3.9k. This has a significant effect on the rise and fall times of the signals, which causes some noise issues. These issues are addressed using digital filtering in the Verilog code. I'll definitely be using real level shifter ICs, if I make some custom boards for the project. The issue could be mitigated also by lowering the resistance values, but that might load the GB output too much.

VGA output from framebuffer

I started with generating the SVGA output from the FPGA. It turns out that generating a VGA signal in Verilog is quite easy. You need two counters to go through the image and blanking area and then assert the sync signals at the proper ranges of the counter values. For more details, see this timetoexplore.net blog entry.

SVGA 800x600 @ 60 Hz operates with a pixel clock frequency of 40MHz, so the system PLL was configured to generate that frequency.

On each positive clock edge
  • The horizontal and vertical counters are incremented, or reset if they would overflow
  • A new framebuffer address corresponding to the new counter values is computed
  • A pixel value fetch from the previous framebuffer address is initiated
  • An output pixel value is asserted from the previous fetch
  • Hsync and Vsync signals corresponding with the active output pixel are asserted
The video signals on the output are two clock cycles behind the counters, as both the fetch address computation and the actual fetching of the data happen in a synchronous manner. The hsync and vsync signals have an additional pipeline synchronize them with the display data.

I would rather have used an asynchronous memory type for the framebuffer, but it seems there is no such thing as an asynchronous 2-port RAM. At least not within the Altera megafunctions available for this FPGA.

First image displayed by my setup. Framebuffer RAM contents are initialized using a .mif file. Only red color channel was connected at this stage.

GameBoy signal decoding to framebuffer

The GB LCD board is connected to the main board with a ribbon cable. The image data is transferred using a seemingly simple protocol over five signal lines: vsync, hsync, clock, data0 and data1.

Looking at the signals with a logic analyzer reveals that the GB sets up the data on data0 and data1 on the rising edge of the clock, and the data is valid on the falling edge. However, it is not entirely clear from just a logic trace how hsync and vsync signals are to be understood.

Many projects online demonstrate the decoding of GB LCD signals into images. However, I've never seen anybody really document the mechanism for the proper decoding of said signals. I, at least, had quote a bit of trouble getting the leftmost column of the image to decode correctly under all circumstances (the GB uses slightly different timings of the signal depending on whether it is displaying sprite data or background tile data).

For the first testing I used a logic capture published by flashingleds.net in their nintendoscope project. This dataset had the problem that I couldn't compare my reconstruction with what the GB LCD was actually displaying. For this reason I later set up my own logic analyzer and captured data from a screen of the level "The Amazon" of the game "DuckTales". This scene has window tiles, background tiles and sprites. It also features a rendering bug in the final pixels of the window. If you're interested in my data, you can download it here.
Correctly decoded image for the data provided by flashingleds.net

Correctly decoded image for the opening in "The Amazon" in "DuckTales".

A rising edge on vsync starts a new frame

After the rising edge, the vsync signal remains high for one scanline of data. There doesn't seem to be anything special in the timing of the falling edge. It is the rising edge, which determines where the new frame starts.

This was determined by guessing the polarity and checking that rows were decoded in the correct order.

A rising edge on hsync starts a new scanline

The GB LCDC will start pushing out data of a new scanline after hsync goes high. More importantly however, the falling edge of hsync latches the data for the first visible pixel on a line.

Clock signal is not valid when hsync is high

There are 160 clock pulses between two hsync rising edges. There are 160 columns in the GB display. This seems like a simple one to one correspondence, but unfortunately this is not the case. One of the clock pulses occurs, while the hsync signal is high. For this pulse, the data is incorrect on both edges. After some investigation and experimentation, it turned out that the correct data for the first pixel is present at the falling edge of hsync.

The data for all other pixels on a line is valid on the falling edge of the clock. The simple approach of sampling all pixels on the falling edge of the clock results in the following decoded images
The image looks strange on the bottom left, but I couldn't be certain if this is how it was supposed to look like.

The first column data decoded on the falling edge of the invalid clock pulse. Notice some black pixels on the left edge, which should not be there. They are a ghost from the treasure chest sprite from the right edge of the previous scanline.
The image from the flashingleds.net data looked weird, but I couldn't be sure if that's how it was supposed to look like. However, after doing some preliminary tests with the decoding logic on the actual FPGA, it became evident that there was something wrong with the way I handled the first column. I had artifacts, where sprites on the right side of the picture were generating ghosts on the leftmost column.

Turning to the DuckTales data, this can be seen on the leftmost column, in which some black pixels are visible where white background should be. These are a ghost image from the last pixel on the previous scanline (the treasure chest sprite seen on the right side of the image). The background and window tile data look correct here, but that is coincidental.

The figure below shows the beginning of the 120th line of the DuckTales capture. This is one of the lines, which contains the treasure chest sprite ghost.
Logic trace of the beginning of the 120th scanline in the capture.
You can clearly see the invalid clockpulse while hsync is high. For the other clock pulses, data is set up on the rising edge, while data is valid on the falling edge. For the invalid pulse, data remains constant on the rising edge while it changes on the falling edge.

The first pixel of this row should be all white, which corresponds with 0 on both data lines. This is clearly not the case near the invalid clockpulse. Most of the time the value decodes as 1 on both lines. This results in a black pixel.

Not quite clear in the figure above, there is a small amount of time of time between the falling hsync edge and the rising edge of the first valid clock pulse / setup of data. During the falling hsync the data is 0, which is what we expect. Based on this observation on this single line, I went on to implement this into the decoding. I can't say with absolute certainty that this is the proper way to decode the signals, but at least I haven't noticed any artifacts pop up. Everything has looked the same as on the the LCD.

Tearing

The VGA output is not synchronized in any way to the GB input. Both are running at approximately 60 frames per second, but not exactly. There is necessarily some beat frequency between the two. I was expecting to see severe tearing artifacts at places, where the GB data input catches up with the VGA output, or where the VGA output overtakes the GB data input. I was pleasantly surprised to not notice any tearing in the output. Due to how the signal is processed, there necessarily is some tearing, but it is much less noticable than I feared.

Media

My GameBoy attached to the FPGA board. Level shift resistors are inside the yellow heat shrink. VGA DAC resistors are inside the HD15 connector.
The whole setup with the VGA monitor I'm using.


A clip showing the whole setup with me playing Super Mario Land

Some more Super Mario Land gameplay now zoomed in to the monitor

Code

The Verilog code can be grabbed here. It uses two Altera megafunction blocks. One ALTPLL block for controlling the PLL and one ALTSYNCRAM block configured as 2-port RAM for the framebuffer.

The whole thing is mostly written in one file and there is no test bench, or any other testing. It is a work in progress, and not complete in any way.

Sunday, July 8, 2018

Frequency counter calibration

In two previous posts I described how I first repaired and then checked the calibration of a Racal-Dana 1992 frequency counter. The oscillator of the counter seemed to be within 100 ppb (parts per billion), which based on the aging specification of the oscillator is about as close as it needs to be calibrated. However, I don't think the unit had been calibrated in a very long time, which hints that the aging specification may be very pessimistic. In any case, calibrating the unit seemed like an interesting thing to do.

The 1PPS output of the GPS receiver I was using to check the calibration is specified at 50ns RMS error. This figure alone isn't quite accurate enough for my needs. However, the mean value of the output should be almost exactly 1 second (ignoring any possible gross error, such as loss of GPS fix). Averaging over a number of measurements should then give a much increased accuracy, which can then be used to tune the oscillator. Averaging 100 measurements should give one tenth of the error, which in this case would be about 5 nanoseconds, assuming a normal distributed error.

To easily get averages over the measurements the easiest route was to use GPIB. Unfortunately I didn't have a GPIB interface on my computer, so I needed to build one. See my previous post on building the interface.

Now that I could get readings from the counter, I wrote a simple piece of software for averaging the counts. The calibration procedure would then be simply:
  1. Acquire 100 measurements
  2. Compute average
  3. Adjust accordingly
After doing this, I collected data over one night (around 7.5 hours) to check the calibration. The results of which is shown below.
PPS period measurement over one night.
There are a few outliers during the night, perhaps from short GPS coverage losses. They do not contribute much to the average, since there are only a few, and the distribution seems to otherwise be very close to the normal distribution.
The mean value of the measured PPS period is just 4.04 +- 0.64 nanoseconds shorter than 1 second, which corresponds with the oscillator running just 4 parts per billion too fast.

The next thing I want to investigate with this setup is the long term stability of the oscillator. For this purpose I'll record the measurements over a long period of time, preferably at least some months. This should give a general understanding of the actual aging behavior of the oscillator, since the specified aging seems to be very pessimistic. This setup should also be accurate enough to capture GPS timing errors like the one observed not too long ago (in Finnish), albeit those are very rare.