Subblock Processing in MMSE-FDE Analysis
Subblock Processing in MMSE-FDE Analysis
Reviewer A Comments:
This paper presents a sub-block processing technique for the FDE implementation for the high
speed channels. In summary, two techniques are used to combat the variations due to the channel:
1. the long data block is divided into small sub-blocks, so that the channel remains relatively
stable within the sub-blocks
2. a pseudo cyclic prefix technique is used to reconstruct the cyclic prefix for each sub-block based
on an iterative approach, so that for each short block, FDE approach is still valid.
My main technical comments are as follows:
A-1
The paper lacks technical depth and novelty. The key technique used in this paper, i.e.
the pseudo cyclic prefix technique is already published in [4], although it is proposed
just to reduce the CP overhead. Here exact same technique is used for each sub-block
instead of the entire block.
A-2
Section II-A essentially derive the FDE implementation for the stationary channel case.
This is based on the well known results that the cyclic matrix can be diagonalized by
the DFT matrix in the form of F HF H = D. It is sufficient just to provide a reference
to this part without repeat the derivation.
A-3
Section II-B derives the condition for fast fading channels. Is there any justification that
the channel can still be formulated as F HF H = D + E, where D is essentially the same
as in the stationary case? Now that the channel is changing with the block of N , strictly
speaking, the definition of D would not be valid any more.
A-4
Section III, need to provide some details for the channel estimation block. Currently,
it only refers to [8]. A short description is here to provide the meaning of “channel
estimation in the time domain”. Also, what is the justification that a cubic interpolation
function should be used to estimate the channel in-between the two pilot bursts?
In summary, it is reasonable that the overall system performance can be improved if the sub-block
processing is enabled by the means of pseudo cyclic prefix. But the current submission does not
have enough technical depth and novelty to justify its publication as a JSAC journal paper.
1
Reviewer B Comments:
Summary: This paper requires mandatory changes to be accepted.
The main contributions of this paper are the derivation of equivalent noise power for MMSE weight
values and the approach that divide a block into several sub-blocks to mitigate the effects of the
time-variation of the channel that corrupts the periodicity.
B-1
The problem of using FDE under fast fading environments is interesting. Such fading
channels destroy the equivalence between time domain and frequency domain represen-
tation of received signals. The same effect is also encountered when OFDM modulation,
which is viewed as the counterpart of FDE, is employed in fast fading channels. The
impact of time variation within a transmission block is well addressed in the following
reference that the authors should consider:
• Stamoulis, S.N. Diggavi, and, N. Al-Dhahir, “Intercarrier Interference in MIMO OFDM,”
IEEE Trans. Signal Processing, vol. 50, no. 10, pp. 2415–2464, Oct. 2002.
When we refer to the similar matrix expressions in the above reference, the block channel
matrices Hb and H given in (4) and (9) in the paper are doubtable. Namely, the time
indices in any row of these matrices should be identical since they all indicate the time
instant of the corresponding received symbol that row represents.
B-2
The authors attempt to model the effect of off-diagonal entries of channel matrix H by
the equivalent Gaussian noise. However, in order to compute the approximated power
for equivalent noise, the authors already relax the fast fading assumption of the channel
model. As shown in the appendix, the approximation from (33) to (34) is only valid
if, for example, h2,−2 equals to h2,N −2 . Of course, this is not true if the fading rate is
very high. Moreover, the authors should provide the theoretical analysis of bit error
probability based on the equivalent noise power to justify the explanation. This can be
easily accomplished since the linear equalization is considered in the paper.
B-3
However, when we consider the simulation result in Fig. 5, this approximation works very
well independently of the number of subblocks - why? (I guess that the approximation
for equivalent noise power is considered only for the first step to obtain the tentative
decisions. Thus, it seems to be independent of the number of subblocks). Moreover, 2-
subblock processing shows best performance for the case of not including equivalent noise
power in Fig. 5. As the size of subblock increases, the accuracy of approximation and
the tentative decision is degraded. It is difficult to understand why this result occurs. In
addition, for the case of introducing equivalent noise power, according to the authors, the
reason why 2-subblock processing shows best performance is that the smaller subblock
size degrades intra-subblock periodicity. Even though we agree that smaller block size
gives a higher ratio of the pseudo CP to the subblock size, the channel within a subblock
seems to be more static with smaller block size. It means that we have some benefits in
terms of FDE operation with smaller size. Thus, the performance gain among different
number of subblocks seems to be not significant as shown in the Fig. 5.
2
B-4
When we compare Figs. 6 and 7, 2-subblock processing shows the best performance for
256 symbols case while 4-subblock processing shows the best for 1024 symbols case. If
the channel characteristics are same for both cases, the subblock processing having same
size seems to show the best performance. But the simulation showed the different result
and it is required to explain more about this phenomenon.
B-5
For the use of FDE in time varying channels, the channel gains at each symbol in the
data block must be estimated. We recall that, under fast fading rate, the channel co-
efficients are not highly correlated. Thus, a specific interpolation technique that takes
into account the Doppler spread should be considered. Going back to the paper, the au-
thors implemented the cubic function to perform the channel estimation. Probably, the
method will produce unexpected results in fast fading channels. It is better to include
sufficient references for the interpolation technique that is used in the paper.
B-6
The subblock processing technique is the natural way to reduce the effect of fast fading.
In this paper, this is done with the introduction of pseudo CP reconstruction. First,
the FDE is applied to the whole data block to obtain tentative decisions of transmitted
symbols. Then, these tentative decisions are used to cancel the ISBI between two adjacent
subblocks and create the pseudo CP. Notably, the equalization scheme that is employed
to obtain tentative estimates is linear equalization. Thus, some errors are expected,
especially in low SNR regime and very high Doppler spread. The subblock processing
needs to remove the ISBI from previous subblock, and reconstruct the pseudo CP based
on the considered subblock. The ISBI cancellation and CP reconstruction, in turn, can
introduce more errors for each subblock processing. This can cause the error propagation
spread among the subblocks. Thus, there must be a trade-off between the number of
sublocks and the error propagation. But, the simulation results in Figs. 6 and 7 do
not reflect this fact. To reduce this, it is better to consider more elaborate equalization
scheme, such as decision feedback equalization. Furthermore, after having the tentative
estimates of transmitted signals, the authors should recognize that the off-diagonal entries
of channel matrix can be eliminated. It is suggested that the authors should perform the
subblock processing after canceling the off-diagonal entries to achieve more performance
improvement.
B-7
It is also required to describe the total process of proposed scheme in detail. For example,
I wonder whether the equivalent noise power is recalculated for the subblock processing
or same equivalent noise power is used for the tentative decision and subblock processing.
B-8
The idea of this paper is interesting. But, it is difficult to verify the contribution with
the current simulation results and discussion about it. The best way is to include the
theoretical analysis. If the theoretical derivation is difficult, at least more sufficient
discussion on the above questions should be mentioned.
Minor comments:
3
B-9
The CP formation and removal are well known, thus there is no need to introduce
the transform matrices for adding CP or removing CP. (equations (26) and (27) are
sufficient).
B-10
The sentence “FD=0.37, which is about 1.5 times the speed of the conventional FDE”
in page 12 should be corrected.
4
Reviewer C Comments:
This paper proposes new frequency domain equalization (FDE) technique which is robust against
fast fading for single-carrier (SC) signal transmission. The received signal block is divided into
shorter subblocks to apply subblock based FDE using decision feedback and pseudo CP generation
technique proposed in [4]. FDE of shorter block size is applied to each subblock using the CSI
obtained by an interpolation technique. Since the channel variation within the data block is tracked
by the interpolation technique, better transmission performance can be achieved. The proposed
FDE seems to be useful.
However, I have a number of comments.
C-1
The proposed subblock based FDE requires the insertion of UW whose length NP is
longer than 2 times the maximum channel time delay L. However, shorter UW length
is desirable from the transmission efficiency view point. For the conventional FDE, the
UW length can be set to NP = L. Of course, the UW length can be made longer while
keeping the transmission efficiency to be the same by extending the FDE block size;
but the tracking ability against fading may be lost. Performance comparison between
the conventional FDE and subblock based FDE should be made assuming the same
transmission efficiency and the same maximum Doppler frequency.
C-2
Sec. II-A is redundant since the FDE operation and the insertion of CP are well known.
C-3
Eq. (21) needs an explanation. Why are s̃i s̃∗j uncorrelated?
C-4
Page 8, 2nd line from the bottom. It is stated that “We need to know the symbol-by-
symbol CSI..”. Why is the symbol-by-symbol CSI necessary? What the FDE requires is
the block-by-block CSI.
C-5
The subblock based FDE requires CSI at each subblock position in the data block of
length N − NP . The 3rd order interpolation technique is used in the paper. However,
no detailed interpolation operation is presented. How is the CSI estimation done for
subblock based FDE within the data block?
C-6
Eqs. (28) to (32) are not necessary since the pseudo CP generation is already presented
using Eqs. (25) to (27).
C-7
Page 12, Sec. V-A. It is stated that “If no error are found in a cyclic redundancy check,
”. This means that the cyclic redundancy check (CRC) code needs to be added to
each data block of length N − NP . This further reduces the transmission efficiency. Or
is the CRC code added to a group of several data blocks?
5
C-8
The normalized Doppler frequency FD is defined as a product of the maximum Doppler
frequency fD and the block length N Ts . If the block length is different, FD is different
for the same fD value. However, in Fig. 8, the BLER is compared for different block
lengths assuming the same FD value (FD = 0.3). Figs. 5 and 6 assume the data block of
256 symbols while Fig. 7 assumes the data block of 1024 symbols, but the same FD = 0.3
is assumed. If the same maximum Doppler frequency fD is assumed, the fading tracking
ability using the interpolation technique should be worse for the data block of 1024
symbols. For fair comparison, the same fD should be assumed instead of the same FD .
6
Authors’ Replies
We wish to express our deep appreciation to the reviewers for their insightful comments on our
paper. All the comments were really helpful to us to improve the paper quality. In the following,
the responses to the comments are listed one by one. For the sake of convenience, we numbered
each comment as follows, and the numbers are provided alongside the corresponding changes in
the revised manuscript. The changes in the manuscript are also shown in this letter. We thank
the reviewers again and hope that the reviewers will find the paper acceptable.
A-1
The paper lacks technical depth and novelty. The key technique used in this paper, i.e.
the pseudo cyclic prefix technique is already published in [4], although it is proposed
just to reduce the CP overhead. Here exact same technique is used for each sub-block
instead of the entire block.
Answer A-1
We agree with the comment that the CP reconstruction technique is not original. Thus, we showed
our respect for the already-published paper [4] by referring to it in the manuscript. However, the
objective of the CP reconstruction in [4] is focused on an improvement of transmission efficiency.
Our key technology in the paper is subblock processing at the receiver side. And, the CP recon-
struction scheme is introduced as a technology to realize our idea. On this point, we believe that
the novelty of the paper is not lost. Indeed, applications of the CP reconstruction to subblock
processing for FDE in fast fading environments have never been mentioned before our proposal.
For technical depth, we have tried our best effort to improve the manuscript quality according
to the comments. The specific changes are shown as the answers for the other comments in the
following.
A-2
Section II-A essentially derive the FDE implementation for the stationary channel case.
This is based on the well known results that the cyclic matrix can be diagonalized by
the DFT matrix in the form of F HF H = D. It is sufficient just to provide a reference
to this part without repeat the derivation.
Answer A-2
According to the comment, we have reduced the descriptions in II-A by referring to a new paper.
However, the derivation part of FDE using the DFT matrix is untouched since we think that it
makes the problems in a fast fading environment more clear. The changes are indicated below.
7
Section II-A
Let us consider a single carrier system with FDE per N -symbol block at the receiver side. To
maintain periodicity of the received signal within a DFT window, we add a CP longer than the
maximum symbol delay. Here, assuming a multipath channel with L symbol-spaced paths, the
CP length NP is set as NP ≥ L. When we define an N -dimensional transmit signal vector
s = [s0 , . . . , sN −1 ]T and an N -dimensional noise vector, the N -dimensional received signal vector
after CP removal is given by [5]
r = Hs + n, (1)
The first index i and the second index j of hi,j correspond to multipath number and transmitting
time, respectively.
Transforming (1) into the frequency domain using the DFT matrix F yields
F r = F Hs + F n (3)
H
= F HF F s + F n, (4)
where F H F = IN because F is a unitary matrix. When the channel is time-invariant within the
block, i.e., hl,0 = · · · = hl,N −1 = hl for l = 0, . . . , L − 1, H becomes a cyclic matrix. Thus, H can
be. . .
Reference
[4] T. Hwang and Y. (G.) Li, “Iterative cyclic prefix reconstruction for coded single-carrier systems with frequency-
domain equalization (SC-FDE),” Proc. VTC2003-Spring, vol. 3, pp. 1841–1845, April 2003.
[5] Stamoulis, S. N. Diggavi and N. Al-Dhahir, “Intercarrier Interference in MIMO OFDM,” IEEE Trans. Sig-
nal Processing, vol. 50, no. 10, pp. 2415–2464, Oct. 2002.
[6] L. Deneire, B. Gyselinckx, and M. Engels, “Training sequence versus cyclic prefix—a new look on single carrier
communication,” IEEE Commun. Lett., vol. 5, no. 7, pp. 292–294, July 2001.
A-3
Section II-B derives the condition for fast fading channels. Is there any justification that
the channel can still be formulated as F HF H = D + E, where D is essentially the same
as in the stationary case? Now that the channel is changing with the block of N , strictly
speaking, the definition of D would not be valid any more.
8
Answer A-3
Thank you for the comment. As in the comment, we should distinguish these matrices. We have
introduced a new matrix D 0 . The changes including the related equations are shown below.
Section II-B
In fast fading environments, cyclicity of the channel matrix H is no longer valid due to channel
transition within the FDE block. Thus, F HF H includes non-diagonal components as
F HF H = D 0 + E, (9)
where D 0 expresses the diagonal component and E is an N ×N matrix having off-diagonal elements
only. Substituting (9) into (4) yields
r̃ = D 0 s̃ + ñ + E s̃ (10)
∑
N −1
r̃k = d0k s̃k + ñk + e(k,j) s̃j , (11)
j=0
where e(k,j) is the (k, j)th element of E. The third terms on the right-hand sides of these equations
correspond to inter-frequency interference components, or may be regarded as residual ISI com-
ponents in the time domain. Thus, the MMSE estimation is only achieved by solving an inverse
problem of the N × N matrix (D 0 + E) so that the numerical complexity grows enormously as the
block size increases.
However, if the third term in (11) is relatively small compared to the first term, we can still
apply the conventional FDE by regarding the third term as additional noise to ñk
A-4
Section III, need to provide some details for the channel estimation block. Currently,
it only refers to [8]. A short description is here to provide the meaning of “channel
estimation in the time domain”. Also, what is the justification that a cubic interpolation
function should be used to estimate the channel in-between the two pilot bursts?
Answer A-4
According to the comment, we have added some descriptions on the channel estimation. For the
cubic interpolation, we use four pilot blocks (two in the past and two in the future of the target
block). So, we have modified the explanation to avoid any misleadings as shown below and added
a new figure (Fig. R-1) showing interpolation to Fig. 1.
9
Section III
where huw is assumed as the time-invariant channel response within the target signal sequence.
Then, huw can be estimated by minimizing J [9] as
We need to know the channel response at the central position of the block/subblock for FDE. In
addition, the approximation of the equivalent noise power estimation requires the channel responses
at the head and tail as in (35) and (42). Moreover, all responses within the pseudo CP part are also
needed. Therefore, the channel estimates at several points are necessary in total. Since the UW
for channel estimation is only located at pre- and post-data blocks, we obtain channel estimates
within the data block by interpolation. In this paper, third order interpolation is used as shown
in Fig. 1(b). By using channel estimates at four UWs (two in the past and two in the future of
the target data block), the channel within the data block is interpolated with a cubic function. To
be specific, a unique cubic curve passing through these four points is solved. Then, the channel
responses at the other required points are obtained from the curve. When applying the FDE. . .
Cubic function
10
B-1
The problem of using FDE under fast fading environments is interesting. Such fading
channels destroy the equivalence between time domain and frequency domain represen-
tation of received signals. The same effect is also encountered when OFDM modulation,
which is viewed as the counterpart of FDE, is employed in fast fading channels. The
impact of time variation within a transmission block is well addressed in the following
reference that the authors should consider:
• Stamoulis, S.N. Diggavi, and, N. Al-Dhahir, “Intercarrier Interference in MIMO OFDM,”
IEEE Trans. Signal Processing, vol. 50, no. 10, pp. 2415–2464, Oct. 2002.
When we refer to the similar matrix expressions in the above reference, the block channel
matrices Hb and H given in (4) and (9) in the paper are doubtable. Namely, the time
indices in any row of these matrices should be identical since they all indicate the time
instant of the corresponding received symbol that row represents.
Answer B-1
Thank you for introducing the valuable paper. First, let us explain the difference in index notation
of the channel matrix element. In the introduced paper, time index is defined by receiving time.
That is, the index is the same as the one of the received signal. In our case, the time index is
defined by transmitting time. That is, the index is the same as the one of the transmitted signal.
Clearly, both notations are equivalent. Because we believe that our notation is useful to explain
the equivalent noise power in the Appendix, we have not changed this notation. We hope that our
representation is acceptable.
Finally, according to the comment, we have reduced the descriptions in II-A with reference to
the new paper. However, the derivation part of FDE using the DFT matrix is untouched since we
think that it makes the problems in a fast fading environment more clear. The specific changes
are shown in the Answer A-2.
B-2
The authors attempt to model the effect of off-diagonal entries of channel matrix H by
the equivalent Gaussian noise. However, in order to compute the approximated power
for equivalent noise, the authors already relax the fast fading assumption of the channel
model. As shown in the appendix, the approximation from (33) to (34) is only valid
if, for example, h2,−2 equals to h2,N −2 . Of course, this is not true if the fading rate is
very high. Moreover, the authors should provide the theoretical analysis of bit error
probability based on the equivalent noise power to justify the explanation. This can be
easily accomplished since the linear equalization is considered in the paper.
Answer B-2
As in the comment, (34) ((24) in the revised paper) equals to (33) ((23) in the revised paper) if, for
example, h2,−2 is the same as h2,N −2 . Unfortunately, the replaced elements have different values
due to fast fading. Thus, this approximation includes some errors as the reviewer thought. (We
11
think any “approximation” includes some errors.) However, the ratio of replaced elements to the
total number of elements in the channel matrix is not large in the assumed single-carrier system.
So, it is expected that the approximation is still valid in our case. Indeed, severe degradation due
to the approximation cannot be seen in Fig. 5.
To support our expectation, we show a correlation chart of the equivalent SNR with and without
the approximation in Fig. R-2. This result indicates that the approximation works well in the low
SNR region, i.e., high impact region for BLER performance. Although some underestimation
cases can be seen in the high SNR region, the effect is not severe as shown in Fig. 5. (It should
be noted that we can reduce the calculation load per frequency point to the order of N from the
order of N 2 by using the approximation in compensation for this small degradation. Thus, total
calculation load decreases to the order of N 2 from the order of N 3 .) We have changed section V-B
as shown below, because we think that showing Fig. R-2 with a short description would be useful
for understanding the approximation validity.
In addition to the above numerical analysis, we tried theoretical analysis. However, the given
period for revision (one month) was not enough to complete it. We would like to continue it as a
future work.
Section V-B
First, we evaluate the validity of the approximation when we calculate the equivalent noise power
as in (17). Figure 5(a) shows a correlation chart of the equivalent SNR with and without the
approximation. It can be seen that the approximation works well in the low SNR region, i.e.,
high impact region for error rate performance. Figure 5(b) shows the block error rate (BLER)
performance with and without the approximation where we recalculate the equivalent noise power
in the subblock FDE. To demonstrate the effectiveness of using the equivalent noise power, the
performance with the noise power only (ignoring the third term in (11)) is also shown. If we do
not use the equivalent noise power, error floors can be seen, and they become worse in the higher
SNR region due to underestimating the noise power. In contrast, the BLER performance improves
considerably by incorporating the equivalent noise power. In addition, very little degradation due
to the approximation is observed. Therefore, in the following, we apply the approximation instead
of the strict calculation.
12
30
20 over estimation
15
10
5
0
−5
Figure R-2 The distribution of the equivalent SNR with and without the approximation when
the block size is 256 symbols and FD = 0.3.
B-3
However, when we consider the simulation result in Fig. 5, this approximation works very
well independently of the number of subblocks - why? (I guess that the approximation
for equivalent noise power is considered only for the first step to obtain the tentative
decisions. Thus, it seems to be independent of the number of subblocks). Moreover, 2-
subblock processing shows best performance for the case of not including equivalent noise
power in Fig. 5. As the size of subblock increases, the accuracy of approximation and
the tentative decision is degraded. It is difficult to understand why this result occurs. In
addition, for the case of introducing equivalent noise power, according to the authors, the
reason why 2-subblock processing shows best performance is that the smaller subblock
size degrades intra-subblock periodicity. Even though we agree that smaller block size
gives a higher ratio of the pseudo CP to the subblock size, the channel within a subblock
seems to be more static with smaller block size. It means that we have some benefits in
terms of FDE operation with smaller size. Thus, the performance gain among different
number of subblocks seems to be not significant as shown in the Fig. 5.
Answer B-3
We are terribly sorry for the lack of description on the subblock processing. The approximation
in the equivalent noise calculation is also used in subblock FDE after dividing received blocks.
Let us discuss a trade-off between subblock size and performance. As the reviewer pointed
out, the channel transition within a subblock becomes more static with a smaller subblock size.
Therefore, the FDE performance would be much more improved if we could reconstruct the pseudo
CP perfectly. Fig. R-3 shows the BLER performance when reconstructed pseudo CP is perfect in
the case of FD = 0.3. From the figure, a monotonic improvement of BLER by decreasing subblock
size can be seen. Comparing this figure with Fig 6(a) gives us a conclusion that reconstructed
pseudo CP is not so accurate due to errors in channel estimation and tentative decisions. Such an
inaccurate pseudo CP adds errors in the received signal and destroys the cyclicity/periodicity of the
13
subblock. Thus, it can be said that the ratio of the pseudo CP to the subblock size highly affects
the capability of subblock FDE. We believe that this is the reason why two-subblock processing is
the best in Fig. 6(a).
We think that showing Fig. R-3 helps the readers understanding of the effect of pseudo CP
accuracy on the BLER performance. So, we have made some changes as below.
Section V-C
The BLER performance is shown in Fig. 6 when the block size is 256 symbols. Figure 6(a)
shows the case that reconstruct pseudo CP is perfect when FD = 0.3. From the figure, a mono-
tonic improvement of BLER by decreasing subblock size can be seen. However, the imperfect
pseudo CP in a realistic situation degrades the performance, and a trade-off between subblock size
and performance is observed. To be specific, the best performance is obtained by two-subblock
processing.
Basically, FDE with a smaller subblock size becomes more tolerant to FD as shown in Fig. 6(a).
However, the performance of the proposed method is affected by the accuracy of the pseudo CP,
which depends on both the channel estimates and the tentative decisions as mentioned above.
Comparing Fig. 6(a) with Fig. 6(b) gives us a conclusion that reconstructed pseudo CP is not so
accurate due to errors in channel estimation and tentative decisions. Such an inaccurate pseudo
CP adds errors in the received signal and destroys the cyclicity/periodicity of the subblock. Thus,
it can be said that the ratio of the pseudo CP to the subblock size highly affects the capability of
subblock FDE. Consequently, two-subblock processing provides the best performance for FD = 0.3.
Figure 6(c) shows the BLER performance versus FD when Eb /N0 = 30 dB. When assuming
a required BLER of 10−2 , two-subblock processing in the estimated CSI case provides the best
performance for FD ≥ 0.3 and is applicable until FD = 0.37, which corresponds to about 1.5 times
the speed applicable in the conventional FDE case, i.e., FD = 0.25.
Next, we show the performance of different block size with the same FD . When FD is the
same in the different block sizes, the channel estimation accuracy is almost the same. So, we can
discuss the relationship between the optimum subblock number and pseudo CP ratio. Figure 7
shows the. . .
14
100
10−1
Average BLER
10−2
no division
2 subblocks
4 subblocks
8 subblocks
10−3
0 10 20 30
Average Eb/N0 [dB]
Figure R-3 BLER performance when CSI and tentative decisions are perfect.
B-4
When we compare Figs. 6 and 7, 2-subblock processing shows the best performance for
256 symbols case while 4-subblock processing shows the best for 1024 symbols case. If
the channel characteristics are same for both cases, the subblock processing having same
size seems to show the best performance. But the simulation showed the different result
and it is required to explain more about this phenomenon.
Answer B-4
In Figs. 6 and 7, the same FD is assumed. Because FD is normalized by the block length, the
maximum Doppler frequency, fD , is different in each case.
When FD is the same, the channel transition within the block is also the same. So, the same
subblock number would give the same performance if the pseudo CP ratio was the same. However,
as mentioned in the previous answer, the FDE performance is affected by the accuracy and ratio
of the pseudo CP. In the same FD case, the channel estimation accuracy is almost the same in
both cases. Since the pseudo CP size is the same in both cases where L = 16, larger subblock
number in the 1024-symbol case shows better performance because the pseudo CP ratio is smaller
than one in the 256-symbol case. (The short discussion is described in the third paragraph of page
12 in the revised manuscript.)
On the other hand, when fD is the same, the channel transition within the block becomes
different in each case. The channel transition within the same subblock size becomes the same
instead. If we could reconstruct the perfect pseudo CP, the same subblock size would show the same
performance. For reference, let us show a typical channel response and the BLER performance in
Fig. R-4 where fD = 0.3/(1024Ts ). (In this case, FD in the 256-symbol and 1024-symbol cases are
0.075 and 0.3, respectively.)
In Fig. R-4(a), the gray solid line shows the actual channel response. It can be seen that the
256-symbol case gives better estimates because the UW insertion is more frequent. Thus, more
accurate pseudo CP is expected in the 256-symbol case. In this case, we have to decrease the
ratio of the pseudo CP in the 1024-symbol case to reduce the effect of the less-accurate pseudo
15
CP. Indeed, as in Fig. R-4(b), the best performance is given by the subblock size of 128 symbols
in the 256-symbol case and the subblock size of 256 symbols in the 1024-symbol case. It can be
said that the accuracy and ratio of pseudo CP are important factors on the FDE performance as
well as the subblock size and the Doppler frequency.
100
Channel response
1
10−1
Average BLER
(real part)
0 10−2
no division
2 subblocks
4 subblocks
256 symbols 8 subblocks
1024 symbols 1024 symbols
actual channel 256 symbols
−1 10−3
0 500 1000 0 10 20 30
Symbol [symbol] Average Eb/N0 [dB]
Figure R-4 Comparison of channel response and BLER performance for two different block sizes
where the maximum Doppler frequency is the same (fD Ts ∼
= 0.0003).
B-5
For the use of FDE in time varying channels, the channel gains at each symbol in the
data block must be estimated. We recall that, under fast fading rate, the channel co-
efficients are not highly correlated. Thus, a specific interpolation technique that takes
into account the Doppler spread should be considered. Going back to the paper, the au-
thors implemented the cubic function to perform the channel estimation. Probably, the
method will produce unexpected results in fast fading channels. It is better to include
sufficient references for the interpolation technique that is used in the paper.
Answer B-5
As the reviewer pointed out, some unexpected results may be seen in the cubic interpolation.
And, it is known that lower-order interpolation is more stable. However, as shown in Fig. R-5,
the best performance is obtained by the cubic interpolation between three interpolation schemes
where FD = 0.3. This means that the first-order and second-order interpolations are not capable
of tracking the channel transition in such fast fading environments. This is the reason why we used
the cubic interpolation in the paper. Because we would like to focus on the subblock processing,
we decided to omit the comparison between these interpolations. We hope that the reviewer agrees
with us.
The cubic interpolation is simple and uniquely determined. Although we did not refer to
references, we have added a short description on it instead. We hope that the changes below are
acceptable for the reviewer.
16
100
10−1
Average BLER
10−2
linear interpolation
2nd interpolation
3rd interpolation
perfect CSI
10−3
0 10 20 30
Average Eb/N0 [dB]
Section III
In this paper, third order interpolation is used as shown in Fig. 1(b). By using four channel
estimates at four UWs (two in the past and two in the future of the target data block), the
channel within the data block is interpolated with a cubic function. To be specific, a unique
cubic curve passing through these four points is solved. Then, the channel responses at the other
required points are obtained from the curve. When applying the FDE. . .
B-6
The subblock processing technique is the natural way to reduce the effect of fast fading.
In this paper, this is done with the introduction of pseudo CP reconstruction. First,
the FDE is applied to the whole data block to obtain tentative decisions of transmitted
symbols. Then, these tentative decisions are used to cancel the ISBI between two adjacent
subblocks and create the pseudo CP. Notably, the equalization scheme that is employed
to obtain tentative estimates is linear equalization. Thus, some errors are expected,
especially in low SNR regime and very high Doppler spread. The subblock processing
needs to remove the ISBI from previous subblock, and reconstruct the pseudo CP based
on the considered subblock. The ISBI cancellation and CP reconstruction, in turn, can
introduce more errors for each subblock processing. This can cause the error propagation
spread among the subblocks. Thus, there must be a trade-off between the number of
sublocks and the error propagation. But, the simulation results in Figs. 6 and 7 do
not reflect this fact. To reduce this, it is better to consider more elaborate equalization
scheme, such as decision feedback equalization. Furthermore, after having the tentative
estimates of transmitted signals, the authors should recognize that the off-diagonal entries
of channel matrix can be eliminated. It is suggested that the authors should perform the
subblock processing after canceling the off-diagonal entries to achieve more performance
improvement.
17
Answer B-6
First, the linear equalization considering the off-diagonal components needs calculation of an N ×N
inverse matrix and also matrix multiplications. These tasks increase the calculation load to N 3
order. This is why we avoided using such high-performance methods as described in Introduction.
The decision feedback to eliminate the off-diagonal components is one of the attractive solutions
as the reviewer pointed out. If the channel estimates and tentative decisions are pefectly known,
it seems to be effective in cancelling the off-diagonal components. However, in such a fast fading
case, the tentative decisions and channel estimates include an unignorable amount of errors as
discussed above. Thus, the performance of the decision feedback type was not so effective as
shown in Fig. R-6. We think that there are some ways to improve its performance. We would like
to continue the study on it as a future work. Thank you for the advice.
100
10−1
Average BLER
10−2
Figure R-6 The BLER performance using the decision feedback to cancel the off-diagonal com-
ponents when the block size is 256 symbols and FD = 0.3.
B-7
It is also required to describe the total process of proposed scheme in detail. For example,
I wonder whether the equivalent noise power is recalculated for the subblock processing
or same equivalent noise power is used for the tentative decision and subblock processing.
Answer B-7
According to the comment, we have added a short description on equivalent noise power calculation
in the subblock processing. The changes are as below.
18
Section V-B
Figure 5(b) shows the block error rate (BLER) performance with and without the approximation
where we recalculate the equivalent noise power in the subblock FDE. To demonstrate the effec-
tiveness of using the equivalent noise power, the performance with the noise power only (ignoring
the third term in (11)) is also shown. If we do not use the equivalent noise power. . .
B-8
The idea of this paper is interesting. But, it is difficult to verify the contribution with
the current simulation results and a discussion about it. The best way is to include
the theoretical analysis. If the theoretical derivation is difficult, at least more sufficient
discussion on the above questions should be mentioned.
Answer B-8
In spite of our effort, currently we cannot complete the theoretical analysis. So, we have added
other numerical results and a discussion based on the helpful comments from the reviewers instead
as shown in this letter. We hope that our revision has improved the technical depth of the paper.
B-9
The CP formation and removal are well known, thus there is no need to introduce
the transform matrices for adding CP or removing CP. (equations (26) and (27) are
sufficient).
Answer B-9
According to the comment, we have deleted the related part.
B-10
The sentence “FD = 0.37, which is about 1.5 times the speed of the conventional FDE”
in page 12 should be corrected.
Answer B-10
According to the comment, we have changed the sentence as below.
19
Section V-C
FD = 0.37, which corresponds to about 1.5 times the speed applicable in the conventional FDE
case, i.e., FD = 0.25.
C-1
The proposed subblock based FDE requires the insertion of UW whose length NP is
longer than 2 times the maximum channel time delay L. However, shorter UW length
is desirable from the transmission efficiency view point. For the conventional FDE, the
UW length can be set to NP = L. Of course, the UW length can be made longer while
keeping the transmission efficiency to be the same by extending the FDE block size;
but the tracking ability against fading may be lost. Performance comparison between
the conventional FDE and subblock based FDE should be made assuming the same
transmission efficiency and the same maximum Doppler frequency.
Answer C-1
When the channel estimation in each FFT block is not required, the UW can be replaced by a
conventional CP (copy of the tail part), the size of which is only L symbols. However, because we
assume fast fading environments, the channel state in each FFT block must be estimated even in
the conventional FDE. Thus, a UW is required for channel estimation for both instead of the CP.
In this paper, we used a simple MMSE channel estimation in the time domain with the UW.
In this case, the genuine UW part not interfered by uncertain data sequences in pre- and post-
data blocks has to be more than L symbols. For example, a block format with UW whose length
is twice the maximum channel delay L has been examined in [8]. Consequently, in both full-
block processing and subblock processing, the UW size becomes the same. In other words, the
transmitted block format is constant irrespective of processing at the receiver as illustrated in
Fig. 1. We hope that the reviewer will agree with us that the assumption is fair.
C-2
Sec. II-A is redundant since the FDE operation and the insertion of CP are well known.
Answer C-2
According to the comment, we have reduced the descriptions in II-A with referring to the new
paper. However, the derivation part of FDE using the DFT matrix is untouched since we think
that it makes the problems in a fast fading environment more clear. The changes are indicated
below.
20
Section II-A
Let us consider a single carrier system with FDE per N -symbol block at the receiver side.
To maintain periodicity of the received signal within a DFT window, we add a CP longer than
the maximum symbol delay. Here, assuming a multipath channel with L symbol-spaced paths,
the CP length NP is set as NP ≥ L. When we define an N -dimensional transmit signal vector
s = [s0 , . . . , sN −1 ]T and an N -dimensional noise vector, the N -dimensional received signal vector
after CP removal is given by [5]
r = Hs + n, (1)
The first index i and the second index j of hi,j correspond to multipath number and transmitting
time, respectively .
Transforming (1) into the frequency domain using the DFT matrix F yields
F r = F Hs + F n (3)
= F HF H F s + F n, (4)
where F H F = IN because F is a unitary matrix. When the channel is time-invariant within the
block, i.e., hl,0 = · · · = hl,N −1 = hl for l = 0, . . . , L − 1, H becomes a cyclic matrix. Thus, H can
be. . .
C-3
Eq. (21) needs an explanation. Why are s̃i s̃∗j uncorrelated?
Answer C-3
Equation (14) corresponds to the autocorrelation function of transmit signals in the frequency do-
main. So, this equation becomes the Fourier transform pair of the transmit signal power spectrum
in the time domain, i.e., a time series of the absolute square of the signal. The transmit signal
power spectrum is constant as shown in Fig. R-7 because we used QPSK and BPSK for data and
UW signals, respectively. Thus, the autocorrelation function of transmit signals in the frequency
domain is given as the Dirac delta function, i.e., the Fourier transform pair of the constant signal.
As shown in Fig. R-8, the autocorrelation function obtained numerically agrees well with this
theory, indeed.
21
P1s
Power
0
0 100 200
Time [symbol]
P1s
Correlation
0 100 200
Frequency
C-4
Page 8, 2nd line from the bottom. It is stated that “We need to know the symbol-by-
symbol CSI..”. Why is the symbol-by-symbol CSI necessary? What the FDE requires is
the block-by-block CSI.
Answer C-4
Thank you for the comment. As the reviewer pointed out, “symbol-by-symbol estimation” might
be excessive. Actually, in the subblock processing, only the channel response at the central position
of the block/subblock is needed for FDE. In addition, however, the approximation of the equivalent
noise power estimation requires the channel responses at the head and tail as in (45) ((35) in the
paper) and (52) ((42) in the paper). Moreover, all responses within the pseudo CP part are also
needed. Therefore, the channel estimates at several points are necessary in total. We added the
description on it. The changes are below.
22
Section III
We need to know the channel response at the central position of the block/subblock for FDE. In
addition, the approximation of the equivalent noise power estimation requires the channel responses
at the head and tail as in (35) and (42). Moreover, all responses within the pseudo CP part are
also needed. Therefore, the channel estimates at several points are necessary in total. Since the
UW for. . .
C-5
The subblock based FDE requires CSI at each subblock position in the data block of
length N − NP . The 3rd order interpolation technique is used in the paper. However,
no detailed interpolation operation is presented. How is the CSI estimation done for
subblock based FDE within the data block?
Answer C-5
We are terribly sorry for the lack of description on the cubic interpolation. Here, we used the
channel estimates at four UW parts and solved a cubic curve passing through all the four points.
We have added the description. The changes are shown below.
Section III
Since the UW for channel estimation is only located at pre- and post-data blocks, we obtain channel
estimates within the data block by interpolation. In this paper, third order interpolation is used
as shown in Fig. 1(b). By using four channel estimates at four UWs (two in the past and two in
the future of the target data block), the channel within the data block is interpolated with a cubic
function. To be specific, a unique cubic curve passing through these four points is solved. Then,
the channel responses at the other required points are obtained from the curve. When applying
the FDE. . .
C-6
Eqs. (28) to (32) are not necessary since the pseudo CP generation is already presented
using Eqs. (25) to (27).
Answer C-6
According to the comment, we have deleted the related part.
23
C-7
Page 12, Sec. V-A. It is stated that “If no error are found in a cyclic redundancy check,
”. This means that the cyclic redundancy check (CRC) code needs to be added to
each data block of length N − NP . This further reduces the transmission efficiency. Or
is the CRC code added to a group of several data blocks?
Answer C-7
In this paper, we consider a data transmission based on automatic repeat request (ARQ). In such
a case, the CRC code is commonly used for a block error check. We hope that the reviewer also
thinks the assumption is reasonable. We have added the assumption in the manuscript as shown
below.
Section V-A
In the following discussions, we use the normalized Doppler frequency FD , which is a product of
the maximum Doppler frequency fD and the block length N Ts (Ts : the symbol duration), as a
fading speed measure.
The receiver configuration is shown in Fig. 4. In this paper, we consider data transmission
based on automatic repeat request (ARQ). Thus, use of a cyclic redundancy check (CRC) code is
assumed to enable a block error check for ARQ. First, the whole received block is equalized and
decoded. If no errors. . .
C-8
The normalized Doppler frequency FD is defined as a product of the maximum Doppler
frequency fD and the block length N Ts . If the block length is different, FD is different
for the same fD value. However, in Fig. 8, the BLER is compared for different block
lengths assuming the same FD value (FD = 0.3). Figs. 5 and 6 assume the data block of
256 symbols while Fig. 7 assumes the data block of 1024 symbols, but the same FD = 0.3
is assumed. If the same maximum Doppler frequency fD is assumed, the fading tracking
ability using the interpolation technique should be worse for the data block of 1024
symbols. For fair comparison, the same fD should be assumed instead of the same FD .
24
Answer C-8
We agree with the reviewer that the same fD is a condition for fair comparison. In this case, the
channel transition within the entire block becomes different. So, we can check the relationship
between the optimum subblock size and pseudo CP accuracy. On the other hand, when FD is the
same, the channel estimation accuracy is almost the same. Then, we can discuss the relationship
between the optimum subblock number and pseudo CP ratio. (The more specific discussion is
written in the answer B-4.) On this point, we think such a comparison may be regarded as fair.
The aim of the comparison is not showing the superiority of 1024-symbol block but discussing the
effect of inaccurate pseudo CP. To avoid misleading, we have made some changes as below.
Section V-C
Next, we show the performance of different block size with the same FD . When FD is the
same in the different block sizes, the channel estimation accuracy is almost the same. So, we can
discuss the relationship between the optimum subblock number and pseudo CP ratio. Figure 7
shows the. . .
25